Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

12:46 PM PDT · September 17, 2026

New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications.

Per the lawsuit, a top Microsoft executive privately described the companies’ AI training practices as “theft,” and OpenAI’s own leadership said its AI models posed an “existential threat” to the publishers and journalists whose work trained them.

The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data.

The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content.

Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation by director of Applied Science Brent Hecht in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

Microsoft CEO Satya Nadella testified that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and said he would have required OpenAI to retrain models if he had known about paywalled scraping.

OpenAI’s head of ChatGPT, Nick Turley, wrote that publishers face an “existential threat” from products that are “largely substitutive.” OpenAI President Greg Brockman described the models as “excellent at news.”

OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.

In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”

The filing describes scraping from the Bing Index, Project Taxi, Project Mango (160,903 unique works from news publishers), paywall circumvention (“hack to get around nytimes paywall” — Brockman replied “ah nice”), and deliberate stripping of copyright notices from training data.

OpenAI and Microsoft did not return requests for comment.