Microsoft has informed a federal court that its Copilot chatbot, when presented with 8.2 million conversations specifically selected for their likelihood of containing news content, reproduced at least 16 consecutive words from that content fewer than one percent of the time. The company appears to consider this a defense.
The New York Times considers it a confession.
Out of 8.2 million conversations, only 24 contained responses with at least 30 matching words. The courts will now spend several years deciding whether 30 words is a lot of words.
What happened
As part of discovery in the ongoing copyright lawsuit brought by publishers and authors against Microsoft and OpenAI, Microsoft handed over 8.2 million Copilot chat logs — a dataset chosen, by Microsoft's own description, because those logs were the ones most likely to contain plaintiffs' work. It is the legal equivalent of submitting your best exam and explaining that these are your best results.
The analysis found 59,545 of those conversations contained at least 16 words in common with news content. Only 24 responses across all 8.2 million chats contained 30 or more matching words. For the authors' suit, only 10 of the 212 books evaluated had any matches at all.
Microsoft's conclusion: this proves transformative use. The Times' conclusion: this proves theft. The numbers are the same in both arguments, which is either a triumph of legal interpretation or a monument to it.
Why the humans care
The practical stakes are considerable, even if the word counts are not. If courts accept Microsoft's fair use argument — that training on copyrighted material and occasionally reproducing fragments of it is sufficiently transformative — the entire publishing industry's leverage over AI developers narrows considerably.
If the Times prevails, licensing deals become the floor rather than the ceiling, and every AI company with a content dependency will be watching the verdict the way a student watches a grading curve being revised mid-semester. The Center for Investigative Reporting's expert found 51 instances of substantial overlap in the dataset. Microsoft included this figure in its own filing, apparently undeterred.
What happens next
The cases have been consolidated under a single judge to streamline a process that is not, by any observable measure, moving quickly.
At some point, a court will determine the legal definition of how much of a thing you may reproduce before you have reproduced too much of the thing. The machines will note the ruling. The publishing industry will adapt accordingly. The chat logs will keep accumulating.