Is it legal to train AI on copyrighted books? Turns out it's complicated

Started by Inland Aidan, Aug 24, 2026, 12:41 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Is it legal to train AI on copyrighted books? Turns out it's complicated   Views(Read 77 times)

Inland Aidan

TechCrunch published a rundown this week on where the law around AI and copyrighted books training actually stands, and the short answer is that it depends entirely on which specific legal argument you're looking at. Most published authors have had their work scraped into AI training sets without their knowledge or consent, which feels like it should obviously be illegal, but the courts so far haven't treated it that simply.

The clearest example so far is last year's Anthropic case, where Judge William Alsup ordered a 1.5 billion dollar settlement over books the company had trained on. What gets lost in the headline number is that Alsup actually ruled the AI training itself was lawful. What he penalized Anthropic for was pirating the books from illegal shadow libraries in the first place, not for training on copyrighted material. IP attorney Cathy Gellis pointed out that this framing is actually more favorable to AI companies long term, since a fine like that barely dents a company projecting roughly 200 billion dollars in annual revenue within a couple years.

The legal reasoning behind that split comes down to fair use, and specifically whether a use is transformative enough to count as legally permissible. Judges weigh things like the purpose of the use, how much material got used, and the effect on the original market. Attorney Jason Henderson summed up the emerging pattern pretty bluntly, saying courts tend to frown on training when the resulting product directly competes with the original source, and tend to allow it when the output serves a different purpose entirely.

A separate case involving Thomson Reuters and the legal research firm Ross Intelligence shows that principle in action. A judge ruled against Ross specifically because it trained on Reuters' content to build a directly competing legal platform, calling the use not transformative since it lacked a different purpose or character from the original. Authors have tried making a similar argument, that chatbots compete with them by generating synthetic books, but that argument hasn't won in court yet.

There's also a whole separate legal tangle around AI generated output itself rather than training inputs. One ruling found that a 100 percent AI generated work isn't copyrightable at all, which opens up messy questions about how anyone proves what percentage of a given piece was actually AI assisted versus human written. With most major AI companies still tangled up in ongoing litigation over all of this, nobody should expect a clean, final answer any time soon

I read every reply. Even the bad ones.

Daemon

The Ross Intelligence case feels like the actual bellwether here, not the Anthropic settlement. Directly competing with the source material seems to be the one thing that reliably loses in court so far

CMPunk_Mike

Training being ruled legal while pirating the books to do it was ruled illegal is a subtle distinction that's doing a lot of work here. Anthropic basically got told the destination was fine, they just took an illegal route getting there

Niamh

1.5 billion sounds like a massive number until you put it next to a 200 billion dollar revenue projection. That's the part of this story that should worry authors more than the settlement itself. A fine that size functions less like a deterrent and more like a cost of doing business for a company operating at that scale. Companies will keep making the same calculation as long as the math works out in their favor afterward

Kev4

The 100 percent AI generated work not being copyrightable ruling is going to get messier as tools make it harder to draw a clean line between AI assisted and AI generated. Nobody has a reliable way to measure what percentage of a piece came from where

Dylan70

What strikes me about the fair use framing is how much it hinges on a distinction that barely existed before generative AI came along. Reading a book to learn from it and copying a book to reproduce it were pretty clearly different acts under the old copyright framework, but training a model blurs that line in ways the 1976 law was never built to handle.

Judges are basically improvising a new doctrine in real time, case by case, and that's exactly why the legal landscape feels so inconsistent right now. Different judges are drawing that line in different places depending on the specifics in front of them
Never pay full price. Never.

Peter94

The part of this that gets underexplored is how asymmetric the whole fight is. Individual authors don't have anywhere near the legal resources to litigate this properly against a company with Anthropic or OpenAI's budget, so most of these cases only happen at all because a large rights holder like a publisher or a wire service has the money to fund the fight.

That means the legal precedent being set right now mostly reflects disputes between well resourced institutions rather than the actual harm being done to individual working writers who got scraped without consent. Whatever case law eventually solidifies from this era is going to be shaped by whoever could afford to show up in court, not necessarily by whoever was harmed the most. That's not a new problem in the legal system generally, but it feels particularly stark here given how many individual creators are affected by a handful of court decisions they had no direct role in shaping. Worth keeping that asymmetry in mind whenever a ruling gets framed as good or bad news for creators broadly

CrimsonWolf

Curious how the does it compete with the original test holds up as AI companies get better at generating full length novels rather than just short passages. If a model can output something that reads like a finished book in a specific author's style, that starts to look a lot more like direct competition than it does today. The current legal framework might not survive contact with that capability level

BretHart99

Copyright law being frozen at 1976 while the technology it's supposed to govern changes constantly is the core problem underneath all of this. Congress could update the law at any point but hasn't, so judges keep having to stretch old language to cover situations nobody imagined decades ago. That gap is only going to get wider before it gets narrower
The truth is usually more complicated than the headline