Is it legal to train AI on copyrighted books? Turns out it's complicated

Started by Inland Aidan, Today at 12:41 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Is it legal to train AI on copyrighted books? Turns out it's complicated   Views(Read 52 times)
Active members in this topic:
Inland Aidan(1)

Inland Aidan

TechCrunch published a rundown this week on where the law around AI and copyrighted books training actually stands, and the short answer is that it depends entirely on which specific legal argument you're looking at. Most published authors have had their work scraped into AI training sets without their knowledge or consent, which feels like it should obviously be illegal, but the courts so far haven't treated it that simply.

The clearest example so far is last year's Anthropic case, where Judge William Alsup ordered a 1.5 billion dollar settlement over books the company had trained on. What gets lost in the headline number is that Alsup actually ruled the AI training itself was lawful. What he penalized Anthropic for was pirating the books from illegal shadow libraries in the first place, not for training on copyrighted material. IP attorney Cathy Gellis pointed out that this framing is actually more favorable to AI companies long term, since a fine like that barely dents a company projecting roughly 200 billion dollars in annual revenue within a couple years.

The legal reasoning behind that split comes down to fair use, and specifically whether a use is transformative enough to count as legally permissible. Judges weigh things like the purpose of the use, how much material got used, and the effect on the original market. Attorney Jason Henderson summed up the emerging pattern pretty bluntly, saying courts tend to frown on training when the resulting product directly competes with the original source, and tend to allow it when the output serves a different purpose entirely.

A separate case involving Thomson Reuters and the legal research firm Ross Intelligence shows that principle in action. A judge ruled against Ross specifically because it trained on Reuters' content to build a directly competing legal platform, calling the use not transformative since it lacked a different purpose or character from the original. Authors have tried making a similar argument, that chatbots compete with them by generating synthetic books, but that argument hasn't won in court yet.

There's also a whole separate legal tangle around AI generated output itself rather than training inputs. One ruling found that a 100 percent AI generated work isn't copyrightable at all, which opens up messy questions about how anyone proves what percentage of a given piece was actually AI assisted versus human written. With most major AI companies still tangled up in ongoing litigation over all of this, nobody should expect a clean, final answer any time soon

I read every reply. Even the bad ones.

Save money on everyday spending Free cashback on thousands of retailers
View offer