AI
The Court Filing Where AI Describes Its Own Appetite
The clearest account of what AI does to news publishing came from inside the AI companies themselves, in emails, internal documents, and sworn testimony unsealed Thursday in the New York Times copyright case against Microsoft and OpenAI. The Times sued in December 2023, alleging that ChatGPT and Copilot trained on millions of its articles and answer user questions with that material. Other newspaper cases were later consolidated into the suit. The newly public filing asks a federal judge to rule for the publishers before trial, and its exhibits read like a guided tour of the industry's own understanding of its appetite.
Start with the Microsoft side. Brent Hecht, the company's director of applied science, wrote in an internal document that the company's AI content strategy had started a cycle that would hurt both model performance and the web at the same time, calling it highly unusual for a product to put pressure on the economic foundations of its essential suppliers. A Microsoft researcher told colleagues that compensating content creators was in the best interest of the employer, the country, and many other groups. Nadella himself testified that anything paywalled should be licensed by anyone who wants to use it for AI training, and said he would have had Microsoft force OpenAI to retrain its models had he known paywalled material was being scraped.
The OpenAI exhibits go further. The head of ChatGPT, Nick Turley, wrote that the products were largely substitutive and would grow more so as they improved, describing the pressure on publishers in plain terms. Internal documents characterized ChatGPT as a modern newsstand. In a presentation, Dario Amodei, then an OpenAI researcher and now Anthropic's chief executive, listed news generation as a top skill of an early model and demonstrated it with the query 'What is the NYT saying today?' After a researcher told OpenAI president Greg Brockman about a way around the Times paywall during scraping, Brockman replied, 'ah nice,' according to the filing.
A separate filing Thursday from authors including John Grisham, George R.R. Martin, David Baldacci, and Jodi Picoult added internal OpenAI discussions about using the book pirating site Library Genesis to train early models. The publishers' twin strategy is now the industry template: sign licensing deals with some AI companies, sue others. News Corp, the Journal's parent, holds content deals with OpenAI and Meta, while two of its subsidiaries have sued Perplexity.
Then came the government's move. The Justice Department filed a statement of interest arguing that training on news content falls within fair use, that the Times' position runs counter to basic copyright principles, and that constraining model development on a misunderstanding of the doctrine would hamper scientific progress and American prosperity. According to Reuters, it was the first time the United States government has taken a position on copyright litigation over AI training data.
The court now holds two pictures side by side: the industry's private candor about where readers get their news, and the government's public case that learning from the news is how progress works. Whichever picture the judge prefers, the exhibits have already changed the conversation, because the most persuasive witnesses for the publishers turned out to be the defendants' own keyboards. For readers, the practical point is brighter than it looks: every licensing deal signed in this fight's shadow is another revenue stream for the people who do the reporting.
Quick answers
What is this story about?
The clearest account of what AI does to news publishing came from inside the AI companies themselves, in emails, internal documents, and sworn testimony unsealed Thursday in the New York Times copyright case against Microsoft and OpenAI. The Times sued in December 2023, alleging that ChatGPT and Copilot trained on millions of its articles and answer user questions with that material. Other newspaper cases were later consolidated into the suit. The newly public filing asks a federal judge to rule for the publishers before trial, and its exhibits read like a guided tour of the industry's own understanding of its appetite.
Why does this story matter?
The court now holds two pictures side by side: the industry's private candor about where readers get their news, and the government's public case that learning from the news is how progress works. Whichever picture the judge prefers, the exhibits have already changed the conversation, because the most persuasive witnesses for the publishers turned out to be the defendants' own keyboards. For readers, the practical point is brighter than it looks: every licensing deal signed in this fight's shadow is another revenue stream for the people who do the reporting.
Sources
- WSJ: AI companies' internal material on publishers
- Reuters: OpenAI and Microsoft executives' quotes
- Reason: DOJ on AI training and fair use
New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.