Microsoft just put numbers on the table in its fight with news publishers and authors, and the company says those numbers work in its favor. According to The Verge AI, new legal filings claim that Copilot almost never reproduces meaningful chunks of copyrighted text, even when users go looking for it. The filings landed Friday as Microsoft pushes for a summary judgment that would end the case early.
This is the copyright battle brought by The New York Times, the Center for Investigative Reporting, book authors, and others against Microsoft and OpenAI. The core accusation: these companies built commercial AI products on copyrighted work and now compete with the people who made it, partly by regurgitating that work back to users.
What the data shows
During discovery, Microsoft handed over 8.2 million Copilot chat logs to an expert hired by news publishers. The Verge AI reports these weren’t random. Microsoft says they were picked specifically because they hit keywords tied to the publishers’ websites, meaning they were the logs most likely to contain infringing material.
Here’s what the analysis found, per Microsoft:
- 59,545 of the 8.2 million chats shared at least 16 words with grounded news content.
- An expert for the Center for Investigative Reporting flagged just 51 instances of “substantial overlap” with CIR’s work.
- In the authors’ suit, only 24 responses contained 30 or more matching words.
- Of 212 books evaluated, just 10 had any matches at all.
Microsoft’s read: even in the logs stacked to find infringement, Copilot rarely spits out full sentences, let alone passages that could replace the original article or book.
Why this matters
This is Microsoft’s fair use argument, made in raw figures. The company’s position is that training AI on copyrighted material serves a “transformative purpose,” and the fact that a model occasionally echoes a line of text “hardly undermines” that. In plain terms: yes, we used your work to build the system, but the system does something fundamentally different from what you do, so it should count as fair use.
What stands out here is the strategy. Instead of arguing law in the abstract, Microsoft is trying to make the infringement claim look statistically tiny. If verbatim copying is as rare as these filings suggest, the “our chatbot substitutes for your journalism” argument gets harder to prove.
The Times isn’t buying it. Lead counsel Ian Crosby called the discovery record damning: “Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry.” He said the publisher looks forward to both companies “being held accountable for their theft.” CIR and the Authors Guild didn’t immediately comment.
The bigger picture
This case is one of the most consequential legal fights in AI right now. The publishers’ and authors’ claims were consolidated under a single judge to streamline things, over the objections of the plaintiffs themselves. And the stakes reach well beyond Microsoft and OpenAI. Nearly every major AI lab trained on scraped web content, so however this fair use question gets answered, it sets the tone for the entire industry.
There’s a political layer too. The Verge AI notes the Trump administration filed a statement of interest in the Times case this week, siding with OpenAI. That’s a signal about where federal sympathies may sit as these disputes play out.
What comes next
Two paths from here:
- If the judge grants summary judgment, the case ends early and Microsoft’s fair use framing gets a major win.
- If the judge sides with the publishers and authors, the fight continues in court, likely toward trial.
For anyone building or deploying AI products, this is worth watching closely. A ruling either way will shape what training data is safe to use, how much verbatim output becomes a legal liability, and whether licensing deals with publishers shift from optional to mandatory. Practitioners should expect more scrutiny on data provenance and output filtering regardless of how this one lands.
The full breakdown of the filings and the numbers behind them is available at the original report from The Verge AI.