technology 6 min read

AI Locusts Are Eating the Linux Kernel Alive

Japanese developers spotted the problem first: AI crawlers are trashing Linux kernel infrastructure with billions of pointless requests, forcing maintainers to shoulder the cost of their own training data. The crisis is only accelerating.

  • AI Training Data
  • Open Source
  • Infrastructure
  • Cyber Security
  • Linux Kernel

The locusts are here, and they’re hungry for your code

Japan sounded the alarm on this one before anyone in Silicon Valley noticed. A Japanese tech column called it a cyber locust plague — 不気味な虫たち, creepy crawlies — and the metaphor is almost too good to make up. Billions of requests pour into git.kernel.org like a swarm descending on rice paddies, consuming what they need and leaving the actual farmers to pick up the tab.

Konstantin Ryabitsev, who runs kernel.org, posted the details on August 29. The numbers are stark: five distributed nodes, 90 CPU cores total, and a flat 20 percent of that processing power — 14 to 16 cores at any given moment — exists solely to render HTML commits for AI scrapers. That is not backup capacity. That is not a temporary spike. That is the baseline cost of doing business when your open-source project becomes someone else’s data mine.

Why the kernel? Why now?

The question most people miss is not why bots are crawling Linux. It is why they are crawling this particular repo, and why they are doing it so violently inefficiently.

Ryabitsev’s answer touches on something deeper than infrastructure: digital prion disease. The term describes a known failure mode in AI development. Train a model on data generated by previous AI models, and the output quality degrades with each generation — a feedback loop of increasingly synthetic, increasingly hollow content. LLM developers know this. They know that feeding models on pre-AI human writing is the only way to buy time before that degradation becomes catastrophic.

The Linux kernel commit history is one of the largest remaining troves of pure, human-authored, pre-AI code on the internet. Every commit traces back to a person with a keyboard and a problem to solve. No AI helped write it. No AI reviewed it. That makes it valuable — and that makes it a target.

But here is what makes the situation grotesque: the AI companies get the data for free, and the Linux kernel project pays for the bandwidth, the compute, and the engineering time required to serve it. The maintainers are literally subsidizing the training pipelines of companies that will monetize the output.

The most stupid way possible

There is a technical detail here that explains why the load is so absurdly disproportionate to the value extracted.

Git repositories are designed for efficiency. You clone once, you get everything. Forks share objects — the same commit never exists twice on disk. The protocol was built precisely to avoid redundancy.

AI crawlers ignore all of this. They do not clone. They make individual HTTP requests asking the server to render each commit as HTML, one at a time. Each request triggers expensive server-side computation. They hit the same commits repeatedly across different forks, generating billions of requests for what amounts to 1.48 million unique commits spread across 922 forks. The math is almost insulting: tens of billions of valid URLs, and the crawlers end up with 922 duplicates.

Ryabitsev’s own description is apt: “the stupidest possible approach — rendering every single commit as HTML and parsing it.”

Daily request volume sits at roughly 6 million. Two-thirds are blocked by Anubis, the proof-of-work challenge system kernel.org deployed to slow bots down. The remaining third — still approximately 2 million requests per day — consists of bots that solved the challenge and got through. By Ryabitsev’s estimate, legitimate human traffic accounts for about 2 percent of total requests. The other 98 percent is machine noise.

The arms race that favors the attacker

What makes this particularly nasty is that the defender always pays to respond, while the attacker can afford to absorb the cost and simply scale up.

Kernel.org’s countermeasures tell the story. First, they blocked obvious bot IPs using fail2ban. The bots switched to browser-identical user agents — Chrome on Windows, Firefox, Safari — and the IP blocks stopped working.

Then the bots started rotating through consumer ISP addresses and mobile networks. They would fire off four or five requests and vanish from the logs entirely. Blocking those IPs was pointless; the attackers had already moved on. “By the time we realized they were bots, the attack was over,” Ryabitsev noted. “They move on to the next target while we’re still recovering, then come back in a different form.”

The attack surface expanded further. The bots now route through proxy SDKs embedded in smart TVs and gaming apps — devices that most homeowners do not know are participating in botnets. “Your TV is probably doing it right now,” Ryabitsev wrote.

The Anubis proof-of-work system was the most creative response. It forces every visitor to solve a computational puzzle — find a SHA-256 hash starting with a certain number of zeros — before accessing the site. For a human, this takes milliseconds. For a bot trying to scrape millions of pages, it multiplies the cost dramatically. Early on, it worked. Bots gave up. Then they adapted. Within months, they were solving difficulty level 4. Kernel.org raised it to level 5, which now causes smartphones to heat up noticeably. The bots are already learning to solve level 5 as well.

This is not a sustainable dynamic. The attacker has infinitely many instances to distribute the computation across. The defender has 90 cores shared with actual users who are trying to browse the kernel tree for real reasons.

What nobody in the West is talking about yet

The Japanese framing of this as a locust plague carries implications that English-language coverage has not fully absorbed. Locusts do not just eat the crop — they consume it faster than it can regrow, and they leave nothing behind for the next cycle.

The Linux kernel project has survived for over three decades precisely because its governance model matches its technical architecture: distributed, meritocratic, and open to anyone who contributes. That openness is what makes it a prize. But it is also what makes it vulnerable to a kind of resource extraction that has no equivalent in traditional infrastructure attacks. This is not a DDoS aimed at taking the site offline. This is a slow, methodical siphoning of compute cycles that everyone benefits from but nobody pays for directly.

The deeper worry is whether open-source projects can adapt their protocols without breaking the very openness that makes them valuable. Ryabitsev has already begun disabling features and limiting crawlable URLs to reduce server load. He has asked users to accept degraded anonymous access as a temporary necessity. The message is clear: the old model is straining under a new kind of pressure.

The hard truth ahead

Ryabitsev’s conclusion is notable for its honesty: “Unclear.” There is no clean answer. New AI model providers appear every day, all hungry for the same clean, pre-AI training data. App developers are finding new revenue streams by turning consumer devices into botnet participants. The economics favor the extractors.

The Linux kernel is not alone in this position. Any large, high-quality, pre-AI open-source repository faces the same exposure. The question Japan’s developers are raising — and one Western outlets have barely touched — is whether the open-source ecosystem needs to redesign its fundamental assumptions about access, or whether it will continue to subsidize the very systems that threaten its long-term survival.

The locusts are not coming. They are already here. And they are eating the rice.