Microsoft Exec Called AI Scraping Theft. Now It’s Unsealed

Threat assessment: high. The New York Times v. OpenAI and Microsoft just got a lot uglier for the defendants.

Newly unredacted filings in the three-year-old copyright case show a senior Microsoft executive privately calling AI training on scraped news “the largest theft of labor in human history,” according to TechCrunch AI. The same filings quote OpenAI leadership admitting its models pose an “existential threat” to the publishers whose work trained them. OpenAI and Microsoft didn’t respond to TechCrunch’s request for comment.

One caveat up front. Most of the new material comes from The Times’ own legal brief, not the underlying exhibits, which stay sealed. The quotes are real, but the context around them isn’t public yet.

What the filings say

Here’s the intelligence, point by point.

  1. Brent Hecht, Microsoft’s director of Applied Science, wrote the “theft” line in a January 2023 memo. He also called it “an astonishing theft of unprecedented proportions.”
  2. A year later, Hecht described a “doom loop.” Microsoft’s own data showed its Copilot answer engine cut click-through rates to nytimes.com by up to 93% compared to regular Bing search. His January 2024 presentation warned this would “hurt the performance of our models and the entire web at the same time.”
  3. Satya Nadella testified that paywalled content should be licensed. Under oath, he said if he’d known OpenAI trained on paywalled material, he would have invoked Microsoft’s right to “require OpenAI to retrain its models.”
  4. OpenAI’s head of ChatGPT, Nick Turley, called the product “largely substitutive” for publishers, and said it “will get more and more substitutive as they get better.” Greg Brockman called the models “excellent at news.”
  5. The paywall workaround. When researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman allegedly replied: “ah nice.”
  6. Copyright notices got stripped on purpose. Researchers “wouldn’t want model outputting” “copyright notices” to users, per the filing.

The scale

The numbers are the part that stands out to me. OpenAI’s mid-training datasets alone hold more than 91,692 copies of works from the NYT, Daily News, and Center for Investigative Reporting. One Common Crawl-derived dataset had over 2 million documents from nytimes.com. A joint Microsoft-OpenAI effort called Project Mango produced a training set with at least 160,903 unique works from the plaintiffs.

The filing also says OpenAI handed Microsoft the entire GPT-3 training dataset, and Microsoft sent data back through Project Taxi and Project Mango. Content also got pulled straight from the Bing Index.

Why this matters

The legal fight turns on fair use. That doctrine lets you use copyrighted work without permission in cases like parody, criticism, or news reporting. Judges have mostly sided with AI companies so far, and the Trump administration filed a brief earlier this month backing OpenAI’s position.

But fair use has a weak spot. One of the four factors asks whether the new use substitutes for the original and damages its market. That’s exactly what these internal documents describe. A 93% traffic drop is market harm in plain numbers. “Largely substitutive” is the defendant’s own product lead using the plaintiff’s favorite word.

Steven Lieberman, counsel for the New York Daily News, put it bluntly in a statement to TechCrunch: “The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong.”

A Microsoft document even flags a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” That’s not a plaintiff’s argument. That’s the defendant’s own slide deck.

What to expect

Three things for practitioners to watch.

  1. Licensing gets more expensive, fast. If Nadella’s own testimony says paywalled content needs a license, every AI lab building on scraped news now has a quote to explain in court. Expect more content deals, and higher price tags.
  2. Training data hygiene becomes a compliance issue. Stripping copyright notices from training data looks bad in front of a jury. Anyone assembling datasets today should assume their pipeline will be discoverable later.
  3. Fair use isn’t settled. Judges have leaned toward AI companies, but no ruling has faced internal admissions like these. This case could be the one that draws the line between transformative and substitutive.

The underlying exhibits remain sealed for now. When they surface, the full context of these quotes will matter a lot. Until then, TechCrunch AI has the detailed breakdown of the filing.

Scroll to Top