Platform Watch
  • GOOGL354.97+1.6%
  • MSFT501.61+1.6%
  • META741.25+11.4%
  • RDDT158.73+5.2%
  • HUBS213.86-1.0%
  • ADBE249.52+0.2%
  • CRM236.42-0.6%
AI Discovery

NYT Filings Reveal OpenAI Knew of Paywall Workaround; Microsoft Executive Called AI Training 'Astonishing Theft'

Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft allege that senior figures at both companies were aware of efforts to access paywalled content and use copyrighted material to train AI models. Internal documents cite a Microsoft executive describing the practice as an "astonishing theft of unprecedented proportions" and internal data suggesting AI-powered search could reduce publisher traffic by up to 93%.

Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft have disclosed what executives and researchers at the companies allegedly knew about the collection of copyrighted news content, including an alleged method for bypassing the Times' paywall.

According to the filings, OpenAI researcher Nick Ryder told co-founder Greg Brockman about a way to access content behind The New York Times' paywall. Brockman allegedly responded, "ah nice." The Times' lawyers argue the exchange shows that Brockman was aware of an effort to access its paywalled material.

The lawsuit, which has been ongoing for three years, centres on whether OpenAI and Microsoft unlawfully used copyrighted material to train AI models. The allegations have not been established as findings by the court, and the underlying exhibits remain partly sealed.

Nadella Testimony and Microsoft's Response

The filings also cite testimony from Microsoft CEO Satya Nadella, who said he would have required OpenAI to retrain its models had he known that paywalled information had been scraped and incorporated into training. Nadella also said paywalled material should be licensed when used for AI training or grounding. Microsoft has said Nadella's remarks should not be interpreted as a legal conclusion on the copyright dispute.

Internal Microsoft Warning

The unsealed material further reveals that Brent Hecht, Microsoft's director of applied science, described the practice of using copyrighted content to train AI systems as "an astonishing theft of unprecedented proportions" and potentially "the largest theft of labor in human history." The filing also says Hecht warned of what he called an "accidental cover up" as efforts were made to identify material from The New York Times and other plaintiffs inside AI systems. Microsoft said Hecht's comments reflected his individual perspective and were not the company's legal position.

Potential Impact on Publisher Traffic

According to internal Microsoft data cited by the Times, Copilot's answer-focused search experience could reduce click-throughs to the Times' website by as much as 93% compared with conventional Bing search. OpenAI executive Nick Turley, who leads ChatGPT, is also quoted in the filing as describing the impact on publishers as an "existential threat," warning that AI products were already capable of replacing some visits to publisher websites and could become more substitutive as the technology developed.

Dataset Allegations

The filings disclose further details about the datasets allegedly containing publisher material. According to the Times, OpenAI's mid-training datasets contained more than 91,692 copies of works from The New York Times, Daily News and the Center for Investigative Reporting. A separate dataset derived from Common Crawl allegedly contained more than two million documents originating from nytimes.com.

The filing also refers to Project Mango and Project Taxi, initiatives through which Microsoft and OpenAI allegedly exchanged or assembled training material. Project Mango allegedly contained copies of at least 160,903 distinct works produced by publishers involved in the litigation. The Times further alleges that researchers removed copyright notices from some training material because they did not want those notices appearing in AI-generated responses.

Neither OpenAI nor Microsoft has responded publicly to the full range of the newly unsealed allegations, beyond the specific statements on Nadella's and Hecht's remarks noted above.

Prepared with AI assistance and reviewed by the editorial team.

Sources

Continue reading

More in AI Discovery →
AI Discovery

AI Search Is Reshaping Retail: What Brands Must Do Now

Conversational AI tools are fragmenting the traditional retail search journey, forcing brands to rethink how they structure product data, create content, and measure visibility. Experts say success now hinges on verified product knowledge graphs, schema markup, and cross-functional teams — not keyword density and backlinks.

7 min read
AI Discovery

Google DeepMind's Dream-RSI Lets AI Agents 'Dream' Their Way to Cheaper, Faster Search

Google and DeepMind researchers have introduced Dream-RSI, a technique that allows AI agents to test alternative search strategies by replaying recorded past attempts rather than running costly new computations. Benchmarked on Gemini 3.1 Pro and Gemini 3.7 Flash across eight tasks, the method matched or outperformed baselines while sharply reducing the number of live attempts required.

3 min read