Copyright and Fair Use Considerations for Text Data in LLM Training

Courts remain split on whether AI training on copyrighted books counts as fair use.

Contributing Editor · · 10 min read
Cover illustration for “Copyright and Fair Use Considerations for Text Data in LLM Training”
Data Licensing · September 23, 2026 · 10 min read · 2,173 words

American publishing pulls in roughly $30 billion a year, with adult fiction and non-fiction alone worth $6.14 billion in 2024. Copyright industries broadly contribute about $2.09 trillion to the annual output of one national economy. GDP annually. That's the economic backdrop against which federal courts spent the first half of 2025 deciding whether feeding copyrighted books, articles, and legal memos into large language models counts as fair use, and the rulings that came back don't agree with each other cleanly enough to settle the question. Transformativeness tends to favor the AI companies. Market harm, and how the training data got sourced in the first place, can just as easily sink them. Anyone building or deploying a language model right now is operating in a legal environment that remains binding and still very much in motion.

Books matter here in a way that's easy to undersell. A model trained on well-edited, professionally structured prose produces noticeably more coherent and accurate output than one trained on scraped web text alone. That's why book corpora became so valuable to AI developers so fast. And it's why so many of them didn't wait for permission. Court filings show Anthropic pulled in at least 5 million books from LibGen and another 2 million from Pirate Library Mirror, both known piracy repositories. That fact alone shapes how courts have started drawing the line between what's defensible and what isn't.

What the four-factor fair use test asks courts to decide

Fair use in one jurisdiction's legal system runs on four factors under a federal statute § 107, and courts are supposed to weigh all four together. No single factor wins the case by itself, though in practice some carry more weight than others depending on the facts.

The first factor, purpose and character of the use, turns mostly on transformativeness: does the new use do something different from the original, or does it just repackage it? Commercial intent matters here too. The Copyright Office's report on AI training singled out this factor, along with the fourth, as carrying "considerable weight" in the AI context specifically.

The second factor, the nature of the copyrighted work, usually favors fair use when the material is factual or already public, and favors the original creator when the work is highly creative, like a novel. Oddly, in the 2025 cases, the works at issue were mostly fiction, the kind of material factor two is supposed to protect, yet courts still treated this factor as a minor player in the overall analysis. The third factor, amount and substantiality, is trickier for AI developers than it sounds: training ingests entire books, cover to cover, which would normally cut against fair use. But the Copyright Office noted that swallowing the whole work isn't fatal if the use is transformative enough, and that technical guardrails preventing a model from spitting out copyrighted text verbatim can soften how much this factor counts against a developer.

Thomson Reuters v. Ross Intelligence: when training replicates the work's market function

Decided in February 2025 by Judge Stephanos Bibas in the District of Delaware, this case gave the clearest example yet of training data doing what it shouldn't: recreating the market the original work served. Ross Intelligence hired a contractor, LegalEase, to build more than 25,000 "Bulk Memos" out of Westlaw's headnotes and its key-number classification system. Those memos then became Ross's training data.

Judge Bibas found the copying commercial and not transformative, full stop. Ross's legal research tool did the same job Westlaw did, for the same customers, using Westlaw's own editorial structure as fuel. On the fourth factor, market harm, the court found Ross's product was a direct substitute for Westlaw, and went a step further: it flagged harm to the emerging market for AI training data itself, not just the market for the original product. That's a theory of harm that reaches past the traditional "does this AI output compete with the book" question and asks whether licensing markets for training data are being undercut too.

Bartz v. Anthropic: fair use for lawfully acquired books, liability for pirated ones

Judge William Alsup, sitting in the Northern District of California, split this case down the middle in a partial summary judgment issued June 23, 2025. The class swelled to nearly 500,000 authors before it was through, though it started with three named plaintiffs, Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, who filed in August 2024.

On legally purchased books, Alsup came down firmly on Anthropic's side, calling the training "quintessentially transformative." Claude doesn't reproduce the books it trained on; it learns patterns of language and style and uses them to generate something new, a process Alsup compared to how a human writer absorbs influence from years of reading. On market harm, he found no sufficient evidence that training copies displaced sales of the original books. His analogy cuts right to it: authors complaining that AI training will flood the market with competing books is "no different than it would be if they complained that training school children to write well would result in an explosion of competing works." But that transformative finding only covered the legally acquired half of Anthropic's corpus. The books pulled from LibGen and PiLiMi weren't protected by the same reasoning, and that sourcing remained a distinct legal problem for Anthropic.

Kadrey v. Meta: transformative training, but a market-dilution theory that could flip future cases

Judge Vince Chhabria, also in the Northern District of California, ruled on June 25, 2025, two days after Alsup, and landed on a similar transformativeness conclusion by a different road. Training Meta's Llama models on books was, in his words, "highly transformative." Books exist to be read start to finish by people; Meta's use pulled statistical patterns and relationships out of the text to power a general-purpose text generator. Different purpose, different character, same conclusion as Bartz on factor one.

Where Chhabria broke from Alsup is procedural, and it matters. Alsup had addressed the piracy sourcing and the training use through distinct analyses, keeping the two questions largely separate. Chhabria refused to slice it that way. He analyzed Meta's downloading of books from shadow libraries and its later use of those copies to train Llama as one single act of reproduction, evaluated together under the fair use test. That's a meaningfully different framework, and it means a future court following Chhabria's approach could weigh piracy directly into the transformativeness analysis rather than treating it as a separate legal problem.

Meta still won on summary judgment, but not because the market-harm question was resolved in its favor for good. The named plaintiffs simply failed to put forward evidence of market harm sufficient to survive the specific record before the court. Chhabria's opinion reads less like a blanket win for AI training and more like a warning: a better-evidenced market-harm argument, brought by different plaintiffs with different data, might come out the other way. The opinion itself noted that the market-dilution theory, the idea that AI-generated works flood a market and suppress demand for the human-authored originals that trained the model, remains open and could carry real weight in future cases.

What the Chakrabarty-Ginsburg-Dhillon study contributes to the market-harm debate

Market dilution has been mostly theoretical in these opinions. A preregistered study by Chakrabarty, Ginsburg (Columbia Law School), and Dhillon (University of Michigan and MIT's Initiative on the Digital Economy), posted to arXiv in March 2026, gives the theory some actual data to stand on.

The design: MFA-trained expert writers and three frontier models, ChatGPT, Claude, and Gemini, were each asked to produce excerpts up to 450 words long, emulating the styles of 50 award-winning authors. Twenty-eight MFA-trained readers and 516 college-educated general readers then judged the results blind, in pairwise comparisons.

Under simple in-context prompting, the results back up what skeptics of AI creative writing have long argued. MFA-trained readers rejected the AI text on both stylistic fidelity (odds ratio 0.16) and overall writing quality (odds ratio 0.13). Even general readers, who showed no particular preference on style (odds ratio 1.06), leaned toward AI on quality (odds ratio 1.82), suggesting untrained readers already have a hard time telling the difference or don't much care.

Fine-tuning flips the entire picture, and this matters for the legal argument. Once ChatGPT was fine-tuned on an individual author's complete body of work, MFA-trained readers reversed course entirely, now favoring the AI output on stylistic fidelity by a wide margin (odds ratio 8.16) and on quality too (odds ratio 1.87). General readers moved even further, favoring AI on stylistic fidelity at odds ratio 16.65 and quality at 5.42. What this shows, empirically, is that fine-tuning on a specific author's corpus can produce output good enough, and stylistically close enough, to plausibly substitute for that author's future work in the market. It's the exact mechanism the market-dilution theory needs to become a live issue rather than a footnote, and it's no longer hypothetical.

The New York Times v. OpenAI and the next phase of litigation

Filed December 27, 2023, in the Southern District of New York, The New York Times Company v. One technology vendor and OpenAI has moved slower and dug deeper than the California cases. As of September 2026, it remains in active pre-trial proceedings with no trial date set, but the procedural fights along the way have been substantial: a contested dispute over ChatGPT user logs.

The most consequential recent development came from outside the courtroom. On September 1, 2026, a federal justice agency filed a Statement of Interest arguing for a training-stage fair use analysis, one that separates the question of training from separate questions about acquisition, storage, and output. The government's position, in plain terms, is that copying works to train a model can itself be fair use, regardless of how thorny the sourcing or output questions get. That's advocacy, not a ruling, and it doesn't bind the court. But it signals where federal policy leans, and it gives OpenAI's side a citable, official articulation of the argument it's been making informally all along.

By September 8, 2026, all three parties, the Times, Microsoft, and OpenAI, had laid out their arguments before the judge. Observers following the case describe it as entering a genuinely new phase, one where the training-versus-output distinction the DOJ pushed for will likely become the central battleground, rather than the more general transformativeness fights that dominated Bartz and Kadrey.

The Copyright Office's May 2025 framework and where it pushes the analysis

Released May 9, 2025, the Copyright Office's 108-page report, Copyright and Artificial Intelligence: Part III, Generative AI Training, is the most detailed federal guidance on this question to date. It doesn't hand AI developers a blank check, and it doesn't hand rights holders one either.

The Office's core position: some AI training can be fair use, but making commercial use of huge troves of copyrighted material to produce expressive output that then competes with the original works in the same markets "goes beyond established fair use boundaries." That's a direct rejection of the idea that training is automatically or inherently transformative just because it's technical in nature. Instead, the Office ties transformativeness to what the model actually does once trained and deployed. A research model that never generates competing commercial content sits in a very different position than a consumer-facing chatbot that generates prose, code, or images competing directly with the works it learned from.

On factors one and four specifically, purpose and character, and market effect, the Office states that these "can be expected to assume considerable weight" going forward, and its own analysis suggests both factors tend to cut against fair use once a model's output starts competing commercially with the training material. That's a notably more skeptical posture than the transformativeness-friendly language coming out of Bartz and Kadrey, and it sets up a real tension between judicial and administrative views of the same technology heading into the next round of litigation.

How German and EU law diverge from that approach, and its consequences for global deployments

U.S. fair use is a flexible, case-by-case doctrine built on judges weighing four factors against a fact pattern. The EU works from a different architecture entirely: a different legal architecture governs text-and-data-mining uses, something with no clean equivalent in that other jurisdiction. equivalent. Other major EU member states follow similarly distinct national frameworks. Rights holders in EU jurisdictions may have avenues unavailable to American authors, who must rely on the open-ended fair use fight now playing out in Delaware and San Francisco courtrooms.

That structural difference has real consequences for anyone deploying a model across both markets. A training approach that clears the bar in the Northern District of California under Bartz's transformativeness reasoning may still run into an enforceable opt-out under EU law if the underlying works came from rightsholders who exercised it. AI companies deploying models across both markets face meaningfully different legal frameworks, one built on judicial balancing, the other on statutory permission structures, and a win in one jurisdiction carries no guarantee of safety in the other.

Filed underData Licensing

More in Data Licensing