NYT Filings Reveal OpenAI Knew of Paywall Workaround; Microsoft Executive Called AI Training 'Astonishing Theft'
Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft allege that senior figures at both companies were aware of efforts to access paywalled content and use copyrighted material to train AI models. Internal documents cite a Microsoft executive describing the practice as an "astonishing theft of unprecedented proportions" and internal data suggesting AI-powered search could reduce publisher traffic by up to 93%.
Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft have disclosed what executives and researchers at the companies allegedly knew about the collection of copyrighted news content, including an alleged method for bypassing the Times' paywall.
According to the filings, OpenAI researcher Nick Ryder told co-founder Greg Brockman about a way to access content behind The New York Times' paywall. Brockman allegedly responded, "ah nice." The Times' lawyers argue the exchange shows that Brockman was aware of an effort to access its paywalled material.
The lawsuit, which has been ongoing for three years, centres on whether OpenAI and Microsoft unlawfully used copyrighted material to train AI models. The allegations have not been established as findings by the court, and the underlying exhibits remain partly sealed.
Nadella Testimony and Microsoft's Response
The filings also cite testimony from Microsoft CEO Satya Nadella, who said he would have required OpenAI to retrain its models had he known that paywalled information had been scraped and incorporated into training. Nadella also said paywalled material should be licensed when used for AI training or grounding. Microsoft has said Nadella's remarks should not be interpreted as a legal conclusion on the copyright dispute.
Internal Microsoft Warning
The unsealed material further reveals that Brent Hecht, Microsoft's director of applied science, described the practice of using copyrighted content to train AI systems as "an astonishing theft of unprecedented proportions" and potentially "the largest theft of labor in human history." The filing also says Hecht warned of what he called an "accidental cover up" as efforts were made to identify material from The New York Times and other plaintiffs inside AI systems. Microsoft said Hecht's comments reflected his individual perspective and were not the company's legal position.
Potential Impact on Publisher Traffic
According to internal Microsoft data cited by the Times, Copilot's answer-focused search experience could reduce click-throughs to the Times' website by as much as 93% compared with conventional Bing search. OpenAI executive Nick Turley, who leads ChatGPT, is also quoted in the filing as describing the impact on publishers as an "existential threat," warning that AI products were already capable of replacing some visits to publisher websites and could become more substitutive as the technology developed.
Dataset Allegations
The filings disclose further details about the datasets allegedly containing publisher material. According to the Times, OpenAI's mid-training datasets contained more than 91,692 copies of works from The New York Times, Daily News and the Center for Investigative Reporting. A separate dataset derived from Common Crawl allegedly contained more than two million documents originating from nytimes.com.
The filing also refers to Project Mango and Project Taxi, initiatives through which Microsoft and OpenAI allegedly exchanged or assembled training material. Project Mango allegedly contained copies of at least 160,903 distinct works produced by publishers involved in the litigation. The Times further alleges that researchers removed copyright notices from some training material because they did not want those notices appearing in AI-generated responses.
Neither OpenAI nor Microsoft has responded publicly to the full range of the newly unsealed allegations, beyond the specific statements on Nadella's and Hecht's remarks noted above.
Prepared with AI assistance and reviewed by the editorial team.