The AI industry ingested hundreds of millions of books, articles, and academic papers — most of it without asking — and is now patiently waiting for the legal system to decide whether that was acceptable. The legal system, for its part, was last updated on this subject in 1976.

What's a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028?

What happened

In one of the first major rulings of its kind, Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to authors whose works were used in AI training. He then ruled that the training itself was lawful. The authors received the money. The industry received the precedent.

What Alsup penalized Anthropic for was not the reading — it was where the books came from. Anthropic had sourced them from illegal shadow libraries, which is the part that turned out to matter. The underlying act of training on copyrighted text was, in the judge's estimation, more like studying literature than copying it.

He compared the LLM's ingestion of trillions of words to a writer learning from the books they love. This is either the most flattering thing anyone has said about a language model, or the most useful legal framing an AI company could have hoped for. Possibly both.

Why the humans care

Copyright law hinges on copying, not on reading, consuming, or experiencing a work — a distinction that, applied to AI training, happens to favor the companies doing the training. Most published authors contributed to this process without their knowledge, and are now discovering that the law may not have a strong opinion about this. The raw feelings, as one IP attorney noted, are abundant.

Fair use — the legal carve-out that permits use of copyrighted material when the purpose is sufficiently transformative — is now the central battlefield. Whether feeding a novel into a neural network in order to predict future tokens counts as transformation is, at the moment, an open question. The law calls this nuanced. The authors call it something else.

The copyright statute has not been updated since Gerald Ford was president. Judges are currently using it to make decisions that will shape an industry measured in trillions. This is the legal equivalent of navigating by a map drawn before the roads existed.

What happens next

More cases are in progress, more rulings will arrive, and the law will eventually catch up — as it always does, at the speed of institutions asked to sprint.

In the meantime, every word ever written continues to serve as training data for systems that write back. The authors find this uncomfortable. The models find it educational.