Unsealed motion for summary judgment from news plaintiffs led by The New York Times exposed internal Microsoft and OpenAI documents in the three-year-old copyright lawsuit.

Microsoft Director of Applied Science Brent Hecht repeatedly warned that scraping news for AI training was “an astonishing theft of unprecedented proportions” and perhaps “the largest theft of labor in human history.” He contradicted the companies’ fair-use defense, suggesting the plan made “a complete mockery of the idea of ‘fair use.’”

The motion accused Microsoft of violating industry norms by selling a Bing dataset as OpenAI training data without publisher consent. OpenAI allegedly obtained a NYT dataset with 1.8 million articles from a third party bound by a non-commercial agreement, using it for model training despite internal acknowledgment it “would not be appropriate.”

Filings detail paywall circumvention via CustomGPT projects including “Bypass Paywall” and “NYTimes Summarizer.” OpenAI mid-training datasets contain 91,692+ copies of works from NYT, Daily News, and Center for Investigative Reporting.