The copyright infringement lawsuit between The New York Times, OpenAI, and Microsoft entered a contentious new phase as unsealed court documents revealed internal debates among tech executives regarding the mass scraping of copyrighted journalism to train artificial intelligence models.
Internal Alarms at Microsoft Over Mass Web Scraping
According to the court filings, senior Microsoft applied sciences director Brent Hecht characterized the large-scale harvesting of web data as “the largest theft of human labor in history,” while internal OpenAI records showed researchers successfully plotting ways to bypass subscriber paywalls.
The newly public records form part of the evidentiary record in the ongoing federal lawsuit, where The New York Times and other media organizations are challenging Microsoft and OpenAI’s reliance on the legal doctrine of fair use.
Disputing Corporate Policy and Traffic Reduction
According to the filings, Brent Hecht warned colleagues in an internal document that the uncompensated ingestion of protected content amounted to unprecedented theft. Microsoft later distanced the corporation from Hecht’s comments, stating through spokespersons that his private remarks did not reflect official company policy or legal analysis.
Additional internal notes uncovered in the litigation show that other Microsoft employees worried the deployment of AI search features would reduce visits to publisher websites.
Microsoft CEO Satya Nadella also addressed the handling of paywalled material during the proceedings. According to his testimony featured in the court documents, Nadella maintained that content protected behind a paywall requires a formal licensing agreement before it can be integrated into AI training sets or used to generate user queries.
Nadella stated that had he known about the ingestion practices at the time, he would have instructed OpenAI to purge the content and retrain the affected models.
Uncovering Workarounds and Subscription Barriers
The unsealed files shed light on technical operations within OpenAI, revealing direct messaging exchanges concerning subscription barriers.

According to the court documents, OpenAI researcher Nick Ryder informed colleagues that he had discovered a workaround to bypass The New York Times paywall. OpenAI President Greg Brockman responded positively to the discovery in the chat logs.
Publishers are leveraging these communications to argue that the developers actively sought unauthorized access to restricted archives rather than relying solely on publicly accessible web crawling.
Weighing Millions of Copied Articles in Federal Court
The exhibits also detail the sheer volume of news articles embedded within training corpora, with plaintiffs pointing to millions of copied articles distributed across multiple development datasets.
The media plaintiffs argue that this direct duplication substitutes for their own subscription services and destabilizes digital publishing revenue models.
District court judges will evaluate these internal disclosures as both defense teams press their summary judgment motions regarding whether generative AI scraping constitutes transformative fair use under United States copyright law.
Keep reading