business 5 min read

How OpenAI Silently Plundered 10 Million News Articles — And Why the World Is Watching

Internal documents reveal OpenAI and Microsoft knew they were harvesting millions of copyrighted news articles without permission — and discussed workarounds for paywalls. The lawsuits underway in New York could redefine the boundary between AI innovation and intellectual theft.

  • Artificial Intelligence
  • OpenAI
  • Microsoft
  • AI Copyright
  • Fair Use
  • News & Media

A theft measured in articles, not dollars

The lawsuit filed by The New York Times against OpenAI and Microsoft sits in a Manhattan courtroom, but its fingerprints stretch across every newsroom on Earth. What is now emerging from court filings and internal documents is not merely a copyright dispute — it is a portrait of an industry that systematically consumed journalism while openly discussing how to sidestep the very paywalls that kept it funded.

According to reports citing court documents, OpenAI scraped over 10 million news articles without authorization to train its GPT models. That number is not an estimate derived from sampling. It comes from Microsoft’s own internal review, conducted by Brent Heath, the company’s chief applied science officer, who described the activity as “an unprecedented scale of astonishing theft” and “the largest labor exploitation in human history.”

The language matters. Heath is not a journalist. He is an executive at the company that funds and distributes OpenAI’s technology. His characterization — delivered internally, not in a press release — signals that even those building the systems understood, at some level, that what they were doing fell outside acceptable practice.

The paywall email that crossed a line

Perhaps the most damning detail in the filings involves a single exchange. One OpenAI researcher reportedly wrote to then-president Greg Brockman describing a method to bypass The New York Times’ paywall. Brockman’s reply, according to court documents: “Ah, that’s nice.”

Read literally, the exchange suggests two things. First, the company had active research into circumventing journalistic pay barriers — not for academic inquiry, but presumably to expand the corpus of content available for training. Second, leadership did not push back. The response was approval, or at minimum acquiescence.

This is the detail that will age poorly in any retrospective of the AI boom. For years, the dominant narrative framed AI development as an unstoppable march of innovation — one that journalists had a duty to cover rather than a right to restrict. The paywall email reframes the story. It shows a company that viewed journalistic paywalls not as legitimate business models but as obstacles to be circumvented.

Why Korean coverage is sharpening the debate

The Korean publication reporting these allegations — Kumkang Ilbo — is notable not only for the detail it is publishing but for the aggressiveness of its framing. English-language outlets such as AFP, the Financial Times, and the Wall Street Journal have covered the same filings, but Korean media has pursued the story with particular intensity. There is a reason for that.

South Korea’s news industry faces the same existential pressure from AI training practices as its American counterparts, but with fewer legal resources and less precedent to rely on. Korean publishers have seen their content absorbed into models like GPT without compensation agreements comparable to those some US outlets are beginning to negotiate. The stories breaking now — about internal awareness, about paywall bypassing, about the sheer volume of unsanctioned scraping — validate fears that have circulated in Seoul for years.

What English readers may not fully grasp is how much of this story is being driven by non-US legal strategies. Korean and Japanese publishers are pursuing parallel claims in their own jurisdictions, building a global pressure campaign that could outlast any single US courtroom outcome.

The fair-use defense — and its limits

OpenAI and Microsoft are mounting a familiar argument: that using copyrighted news articles for AI training constitutes fair use because it is transformative. The US Department of Justice has filed a brief supporting this position, warning that narrowing the fair-use doctrine would harm competition in the AI market.

But the fair-use argument rests on a premise that the latest documents challenge. Fair use traditionally assumes good-faith engagement with existing law — not the systematic circumvention of paywalls designed to monetize the very content being consumed. If the court accepts the paywall-bypassing emails as evidence of willful infringement, the fair-use defense loses much of its moral and legal force, regardless of what the DOJ argues.

There is also a practical dimension. A ruling that broadly protects AI training under fair use effectively transfers the economic value of journalism from news organizations to technology companies — with no mechanism for redistribution. That is a policy outcome, not just a legal one, and courts are increasingly aware of the difference.

Who wins, who loses, and what comes next

If OpenAI and Microsoft lose in New York, the damages could be measured in billions — but the real cost would be structural. Every AI company training on copyrighted text would face the same exposure. The entire industry model, built on the assumption that scraping is free, would require renegotiation.

If they win, the precedent would lock in the current arrangement: journalism as unpaid raw material for a multi-trillion-dollar industry. That outcome would accelerate the collapse of subscription-based news models, since AI systems trained on copyrighted content can replicate summaries, analyses, and reporting — exactly the products subscribers pay for.

The timeline matters, too. The lawsuit was filed three years after the alleged scraping occurred. OpenAI and Microsoft have had years to adjust their data practices. Yet internal communications suggest they continued operating under the same assumptions well into the training process. That persistence undermines any claim that the company later acted in good faith once alerted to the problem.

The broader reckoning

What is unfolding in this case is not simply a copyright dispute between a newspaper and a tech company. It is the first major trial of whether the AI industry’s foundational practice — scraping the internet for training data — can survive legal scrutiny when that internet includes journalism funded by subscriptions.

The 10 million figure is a benchmark, not a ceiling. Other publishers worldwide are likely to file similar claims, each contributing to a legal architecture that could redefine what constitutes fair use in the age of artificial intelligence. Korean publishers, already active in this space, will be watching closely.

The outcome will determine whether newsrooms remain creditors in the AI economy or simply its unpaid suppliers. No amount of fair-use rhetoric changes that fundamental question.