Every major AI lab has built its models on an enormous diet of text — and increasingly, courts are being asked whether that diet was legally obtained. Over the past year, the question “did you have the right to train on this?” has gone from an academic debate to a multi-billion-dollar legal reckoning, with new lawsuits landing almost monthly.
Here’s where things actually stand.

The books came from somewhere — and that’s the problem
Reports this year detailed how some AI labs, including Anthropic, sourced training material by buying physical copies of books, cutting off the bindings, and scanning the pages — a workaround to build a legally purchased digital library. That approach sits in sharp contrast to another method that’s drawn far more legal heat: pulling books directly from pirated datasets and shadow libraries like Books3, Library Genesis, Z-Library, and Anna’s Archive.
That distinction — purchased-and-scanned versus pirated-and-downloaded — has become the fault line running through nearly every case in the courts right now.
The Anthropic settlement: a number the whole industry is watching
The clearest signal of where this is headed came when Anthropic agreed to a $1.5 billion settlement, resolving claims that it trained Claude on pirated copies of books. Under the deal, thousands of authors are set to receive roughly $3,000 per book. It’s the largest resolution of its kind so far, and it effectively puts a price tag on unauthorized use of pirated text — a number every other AI company now has to weigh against the cost of just licensing content properly.
Notably, an earlier ruling in the same case had actually sided with Anthropic on one core question: training on legally purchased books can qualify as fair use. What settled the case wasn’t the training itself — it was the piracy used to obtain some of the material in the first place.
Meta’s fight: did leadership personally sign off?
Publishers including Hachette, Macmillan, McGraw Hill, Elsevier, and Cengage — joined by novelist Scott Turow — filed a class action against Meta and Mark Zuckerberg personally, alleging Meta trained its Llama models on millions of pirated books and journal articles pulled from the same shadow libraries. The complaint doesn’t stop at the company; it alleges Zuckerberg himself authorized the approach, a claim that, if it holds up, could reshape how personally liable tech executives are for their companies’ training data decisions.
Google’s newest lawsuit: when a trusted program becomes a legal liability
In July, a coalition of publishers — Hachette, Cengage, Elsevier, and author Scott Turow among them — sued Google, alleging it trained Gemini on books obtained through programs like Google Books and titles uploaded to the Google Play Store. The twist here is relational: publishers had handed those books over for narrow purposes, like enabling searchable snippets, not full-scale AI training. The lawsuit reportedly cites an internal Google document acknowledging that using those books for AI training could be “highly problematic” and expose the company to fines in the tens of billions.
That detail matters. It suggests this isn’t simply a case of AI labs being caught off guard by an ambiguous legal question — some appear to have understood the risk internally before proceeding anyway.
It’s not just books anymore
The same legal logic is spilling into adjacent industries. Music publishers have sued Anthropic directly over Claude’s ability to reproduce copyrighted song lyrics on request — testing a narrower and arguably more dangerous question for AI companies: even if training itself is fair use, is a model that can spit out the original material on demand a separate infringement? News organizations are in the fight too — CNN has sued Perplexity, alleging the company copied thousands of articles to train systems that then generated text nearly identical to CNN’s own reporting.
The legal theory courts keep coming back to
Almost every one of these cases runs through the same four-factor fair use analysis, but one factor has emerged as the real battleground: does the AI’s output compete with the market for the original work? An earlier landmark ruling against Thomson Reuters’ AI-tool rival Ross Intelligence set the tone — the court found that training a legal research tool on Westlaw’s original content wasn’t fair use specifically because the resulting product competed directly with Westlaw. That “market harm” framing is now the lens through which judges are evaluating nearly every subsequent case.
Where this leaves AI companies — and everyone else
The emerging split is becoming clear: training on properly licensed or legally purchased material stands a real chance of holding up as fair use. Training on pirated material does not — and it’s already proven expensive enough to force a $1.5 billion settlement rather than risk a trial. That’s likely why licensing deals, not lawsuits, are quietly becoming the industry’s preferred way forward: some AI-training licenses have reportedly been offered at rates around $10 per title, a sign that companies increasingly see paying for content as cheaper than litigating over it.
The definitive answers are still coming. Major trials, including The New York Times v. OpenAI and Getty Images v. Stability AI, are expected to play out over the next year and will likely set precedent far beyond the companies actually named in them. Until then, every AI lab training a new model is making a bet: that its sourcing will hold up if a court ever asks the same question Anthropic, Meta, and Google are now being forced to answer.
