They Knew the Web Was a Doom Loop. They Built It Anyway.
Unsealed court documents reveal OpenAI and Microsoft internally warned that their AI scraping would trigger a 'doom loop' for web content — a cycle they proceeded to accelerate. What their own words expose about the cost of AI growth.
The Documents That Should Have Raised Alarms
Unsealed court filings in The New York Times’ lawsuit against OpenAI and Microsoft have produced one of the most candid admissions in the history of AI business strategy. The companies’ own internal documentation warned that their approach to data scraping would start a “doom loop” damaging the web. Their researchers characterized the scale of content harvesting as the “largest theft of labor in human history.” Their legal positioning made a “complete mockery of the idea of fair use.”
These warnings came from inside the company. They were not leaked whistleblowers or hostile experts. They were written by people employed to build the very systems that are now reshaping the internet.
The most prominent voice belongs to Brent Hecht, Microsoft’s Director of Applied Science. His characterizations are blunt by any standard, but they are not naive. Hecht understood what was happening because he was helping design it. And now his company is trying to write him out of the record.
Microsoft’s Damage Control
Microsoft’s response to the published quotes has been swift and carefully constrained. Spokesperson Alex Haurek told The Verge that the comments reflect “one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.”
In a separate filing, Jordan Usdan, Microsoft’s GM for Data Strategy and Ops, went further — characterizing Hecht’s role as adversarial. He described Hecht as someone who “holds divergent, academic, and forward-looking views about how data ecosystems for AI should operate” and noted that Microsoft employs him precisely to bring “asymmetrical, futuristic, and academic points of view.”
The framing is telling. Microsoft is not denying the content of Hecht’s analysis. It is attempting to contain his authority — reducing a structural warning to the personality of a single researcher whose views are, by design, unconventional within the company.
But the damage of that distinction is limited. The doom loop was not a fringe theory. It was a documented internal conclusion. And whether Microsoft wants to own Hecht’s specific phrasing, the outcome he predicted is now visibly unfolding.
Google Zero Is Not a Projection. It’s Real.
The doom loop Hecht described is no longer hypothetical. Google Zero — the phenomenon where Google’s search results increasingly return AI-generated summaries drawn from scraped content rather than links to original sources — is live and measurable. When AI models consume web content at scale, they incentivize publishers to gate their material behind paywalls or remove it entirely. As original content disappears from the open web, AI models have less raw material to learn from. The models degrade. The platforms respond by scraping even more aggressively from the remaining open sources. Publishers retreat further. The loop tightens.
This is the doom loop. It is self-reinforcing. And it is happening now.
The New York Times is not the only publisher noticing. Outlets across the industry have reported declining referral traffic from search engines, altered caching behavior, and active efforts to block AI crawlers. Some have succeeded. Most have not. The economics of the open web were already fragile. AI model training has accelerated the pressure.
Who Pays for the Loop
The doom loop does not distribute its costs evenly. Publishers and creators bear the immediate burden — lost ad revenue, reduced organic traffic, diminished ability to monetize their own work. Newsrooms that funded investigative reporting through digital subscriptions and display advertising are seeing those revenue streams compress as AI intermediaries capture attention without contributing to the content ecosystem.
The bigger question is who profits from the loop. OpenAI and Microsoft have built models that would not exist at current capability levels without access to decades of human-created web content. The margin on that content, from the perspective of the companies building the models, has been effectively zero. The licensing deals they have struck — including the widely reported New York Times agreement — represent a partial correction, but they are deals with the most prominent publishers. Smaller outlets, independent journalists, and niche creators receive neither compensation nor negotiation leverage.
The asymmetry is structural. A single newspaper can refuse to license its content. An AI company can scrape it anyway and adjust. The legal uncertainty favors the platform.
What Nadella and Altman Say Matters Less Than What They Did
The court documents also reference statements from Satya Nadella and Sam Altman. Neither man is immune to the contradiction at the center of this situation: public rhetoric about responsible AI coexists with internal documentation warning that the current data acquisition strategy is unsustainable and potentially exploitative.
That contradiction is not unique to OpenAI or Microsoft. It is the defining feature of an industry that built its valuation narrative on the promise of abundance while quietly preparing for scarcity. The models require ever more data. The data source — the open web — is being degraded by the same process that consumes it. The companies are aware of this dynamic. Their filings prove it.
The Real Stakes
The Unsealed documents matter because they strip away the abstraction. This is not a debate about fair use doctrine or the philosophical question of whether text-and-data mining is inherently beneficial. It is a record of a company understanding its own behavior, documenting the consequences, and continuing regardless.
The question for courts, regulators, and the public is not whether the doom loop is real. It is whether the companies that predicted it should be held accountable for accelerating it.
The answer to that question will determine the architecture of the web for the next decade. If the current trajectory continues, the open web becomes a resource extraction zone — a place where content is mined until it is no longer viable to produce, then abandoned. If the opposite path is chosen, the incentives shift. Publishers gain leverage. Creators gain compensation. The loop breaks.
Microsoft and OpenAI knew which direction they were heading. The documents make that clear. The legal cases now underway will decide whether that knowledge matters.
What Happens Next
Several outcomes are plausible. The most likely is a settlement framework that extracts licensing payments from publishers without establishing clear precedent for future cases. This would provide temporary relief for major news organizations while leaving smaller creators exposed. The doom loop would continue, albeit at a slightly reduced pace.
A less likely but more consequential outcome is a judicial ruling that establishes binding limits on AI data scraping. Such a decision would reshape the economics of the entire industry. Models that depend on unrestricted web access would face either higher training costs or capability reductions. The market would respond accordingly — some companies would survive, others would not.
The third possibility is regulatory intervention at the legislative level, potentially in the European Union under the Digital Services Act framework or in the United States through new copyright legislation. This path is slower but potentially more durable than litigation alone.
Regardless of which scenario unfolds, the internal documents are now part of the public record. They provide evidence that the companies involved understood the consequences of their data strategy. That knowledge will matter in court. It may also matter in the court of public opinion, where the narrative around AI development is still being written.