The $1.5 Billion Copyright Settlement That Left Japanese Authors With Nothing
Anthropic agreed to pay $1.5 billion to settle a copyright lawsuit over pirated books used to train Claude. But the settlement largely excluded non-English creators — including Japan’s Haruki Murakami — revealing how AI licensing structurally disadvantages creators outside the Anglophone world.
The Settlement That Skipped Half the World
Anthropic agreed to pay $1.5 billion to settle a copyright lawsuit. The money compensates authors whose books were downloaded from pirate sites and used to train Claude, the company’s large language model. On paper, it looks like a landmark victory for creators’ rights in the age of artificial intelligence.
But if you look past the headline number, the settlement tells a different story — one where Japanese authors like Haruki Murakami received virtually nothing, despite their work being among the data Anthropic scraped.
This is not an accident. It is the structural outcome of an AI licensing system that treats English-language rights as the default and everything else as secondary.
How the Lawsuit Began
The case started in August 2024, when three authors — Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson — filed a class-action lawsuit in the U.S. District Court for the Northern District of California. Their claim was straightforward: Anthropic had systematically downloaded approximately seven million books from pirate databases LibGen and PiLiMi without permission, then used them to train Claude.
The case number, 4:24-cv-05417, became a reference point for everyone watching how American courts would handle AI and copyright. By June 23, 2025, Judge William Alsup had issued what he called his “Order on Fair Use” — Document 231, a 40-page ruling that drew a careful line.
Alsup concluded that using books to train large language models was “highly transformative.” AI does not reproduce novels or replace them; it learns patterns, structures, and techniques to generate something entirely new. That reasoning aligned with the core fair-use doctrine in U.S. copyright law, which prioritizes whether a use adds new expression or meaning rather than merely copying the original.
But Alsup drew a sharp distinction on a second point. Keeping pirate copies as a permanent archive was not fair use. “Creating a general-purpose permanent library is not fair use that excuses piracy,” he wrote. That finding created the pressure that led to settlement negotiations.
The $1.5 Billion Number
In August 2025, Anthropic and the plaintiffs reached a basic settlement agreement. The total: $1.5 billion. Alsup gave preliminary approval in September 2025. On July 20, 2026, Judge Araceli Martinez-Olguin entered final approval.
The money compensates for two things: the unauthorized download of books from pirate sites, and Anthropic’s decision to retain those copies rather than destroy them. It does not compensate for the training itself — Alsup had already ruled that part lawful.
So what does $1.5 billion buy? Essentially, it buys Anthropic’s freedom to keep the data it already has. The settlement is a price on persistence, not on use.
The Japanese Gap
Here is where the story widens beyond American copyright law.
The pirate databases Anthropic scraped — LibGen and PiLiMi — contain books in dozens of languages. Japanese titles are heavily represented. Murakami’s works, including “Norwegian Wood,” appear in these archives. Anthropic’s scraping tools would have pulled them in alongside English-language titles.
Yet the settlement produced almost no payout to Japanese creators.
The reasons are structural, not accidental.
First, the class action was filed under U.S. law by U.S. plaintiffs. The named authors — Bartz, Graeber, and Johnson — are American. The legal mechanism for compensation flows through American copyright statutes and American court procedure. Japanese authors holding rights in Japan have no automatic standing in this case.
Second, Anthropic’s settlement pool was calibrated to the works of the named plaintiffs and similar American authors. The company negotiated with the class counsel, not with foreign rights holders. There was no outreach to the Japan Publishers Association, no partnership with rights collectives like the Society of Authors, no mechanism for a Japanese estate to claim a share.
Third, and perhaps most importantly, the settlement’s economics reflect the commercial value of English-language catalogs. Seven million books were scraped. The vast majority are in English. The settlement amount tracks the licensing value of that English corpus, not the global multilingual corpus.
A Japanese author whose work appeared in a pirate database has no pathway into this payout. The system simply does not recognize them as a party.
What This Reveals About AI Licensing
The Anthropic settlement is being hailed as a milestone for creator compensation. It is also a milestone for something else: the exclusion of non-Anglophone creators from AI licensing deals.
This pattern will repeat. Every major AI company is scraping multilingual data. Every settlement will follow the same legal architecture — U.S. class actions, U.S. statutes, U.S. plaintiff pools. The money will flow to the authors named in those cases. Everyone else watches.
The gap is not a bug. It is a feature of how intellectual property law intersects with AI development. Copyright is territorial. AI is global. The mismatch produces winners and losers along linguistic lines.
English-language creators sit inside the legal mechanisms. Japanese, Korean, French, and German creators sit outside them — even when their work is equally represented in the training data.
Why It Matters Globally
Japan is the third-largest book market in the world. Murakami is one of the most translated living authors. His estate should not be invisible in a settlement about the books AI companies learned from.
The same exclusion applies to Korean web novels, German nonfiction, Chinese literature. All of it ended up in pirate databases. All of it trained models. None of it reached the compensation pool.
This is not a dispute about whether AI should learn from books. The courts have already ruled on that. This is a dispute about who gets paid when the learning happens — and who gets left out because their rights exist in a jurisdiction the settlement does not recognize.
The $1.5 billion figure will be quoted for years. What it does not show is the hundreds of millions of dollars in value that flowed to AI companies from non-English catalogs with zero compensation to the creators who produced it.
That number is harder to calculate. It is also the more honest measure of what this settlement actually represents.
The Path Forward
A few things could change this dynamic.
Rights collectives in Japan, Korea, and Europe could negotiate directly with AI companies, bypassing U.S. court mechanisms entirely. Multilateral licensing frameworks could recognize training data as a category requiring cross-border compensation. AI companies could build transparency into their scraping practices and voluntarily include non-English catalogs in settlement pools.
None of these are guaranteed. The current architecture favors the path of least resistance: settle in U.S. courts, pay English-speaking plaintiffs, move on.
For Japanese authors and their estates, that path ends with a press release about a $1.5 billion settlement and a silent empty envelope.
The data was used. The models were trained. The creators were not consulted. And the check never arrived.
That is the story the headline number does not tell.