Memorandum
- From
- Connor Quincy via Fast Company
- Date
- Filed
- Business·5 min to read
- Re
AI Companies Buy Publisher Content but Leave Pricing Power in Doubt
ReAI Companies Buy Publisher Content but Leave Pricing Power in Doubt
Court documents and market data show AI firms recognize the value of publisher content, yet a functioning market for real-time inference data remains elusive as scraping outpaces licensing.
Artificial intelligence companies have begun paying publishers for content, but the emerging market remains one in which sellers have little influence over price. Internal documents made public in The New York Times's copyright lawsuit against OpenAI and Microsoft reveal that executives at both companies understood the value of the material their systems were consuming, even as they moved to take it.
Microsoft applied science director Brent Hecht warned in a memo in early 2023 that millions of people would soon view large models as «hoovering up» their work, calling it «the largest theft of labor in human history.» When a researcher described getting around the Times paywall, OpenAI President Greg Brockman replied, «ah nice.» Nick Turley, OpenAI's head of ChatGPT, described chatbots as an «existential threat» to publishers. The filings underscore a point that has become central to the legal fight: content has value, even when it is scraped from the open web for free.
The Department of Justice weighed in on the fair-use question, filing a statement of interest in the Times case. The DOJ argued that training AI models on publishers' content qualifies as fair use, a position consistent with the administration's view that concessions on copyright would disadvantage the United States in its AI competition with China. President Donald Trump summarized the argument bluntly: «China's not doing it.»
That position addresses training, not inference — the process by which AI systems generate answers about current events. The DOJ acknowledged that «an output reconstructing and disseminating an original copyrighted work may not be transformative,» a description that closely fits AI search, which summarizes current reporting. The unsealed documents also speak to another pillar of fair-use doctrine: harm to the market for the original work. Turley wrote that OpenAI's products «are largely substitutive, period,» and Microsoft CEO Satya Nadella testified that using chatbots «has substituted» for visiting original sources.
The Times filed its lawsuit in December 2023, when the debate centered on training data. That ambiguity has made it difficult for a market to form around training data, since it is hard to justify investing in a payment framework when indicators suggest the material might be free in a few months. Training is where the lawsuits are; inference is where the money is.
Providing AI answers about real-time content is the more valuable business, and AI companies should logically be lining up to pay for the quality content that improves their products. That has largely not happened. Despite scattered licensing deals, the industry has shown little interest in building a marketplace or payment technology to buy content in real time at a fair price. Brian Morrissey of the Rebooting suggested most content is not unique enough to command payment — even the world's best enchilada recipe is only needed once. Commoditization drives a race to the bottom, and that bottom is close to zero.
The broader market tells a different story. Data brokers that scrape the internet at scale and resell the data — firms such as Exa, Parallel, and Tavily — serve not only AI companies but also ad agencies, investment firms, and even other publishers. One estimate cited in Matthew Scott Goldstein's report on the scraper economy put that market at about $1 billion. A market for inference exists, but the money mostly bypasses the media.
The problem is worsening. Bot-protection company DataDome reported that «bad» bot traffic grew 124% in a year, more than nine times faster than human traffic, meaning more bots visit sites and scrape data even when told not to. Scraping alone jumped 185%. Meta, which released its personal AI agent Muse, accounted for 46.3% of all AI bot traffic DataDome tracked in the first half of the year, well ahead of OpenAI. Meta also has deals with publishers including USA Today, CNN, Fox News, and People Inc. The AI world's biggest data harvester is also a buyer — it simply decides when to pay.
Redirecting the inference market back toward publishers will take work, but it is beginning. Parallel has introduced a way to pay publishers for their contributions to agent responses, a sign that some intermediaries see a path to compensating the sources their systems rely on.
2
