Google Freezes Bug Bounty as AI Hallucinations Overwhelm Open Source
Google has paused its open-source bug bounty program until at least Q1 2027, citing a flood of AI-generated submissions that are mostly invalid. The move signals a growing crisis at the intersection of generative AI and software maintenance.
Google’s open-source bug bounty is frozen. Here’s why that matters.
Google has paused its Open Source Software Vulnerability Rewards Program until at least the first quarter of 2027. The official reason: a significant rise in automated submissions, most of them invalid or hallucinated. The freeze goes into effect October 1, and while Google has directed researchers toward its other bug bounty programs, those cover narrower scopes — commercial products, not the sprawling ecosystem of Google’s own open-source projects.
This is the first major platform to formally quarantine what the industry has been calling AI slop. The decision reveals something uncomfortable about where generative AI is heading. When models can generate convincing but fabricated security reports at scale, they don’t just clutter inboxes. They undermine the infrastructure that keeps software safe.
What happened here
According to posts on X and the program website, Google engineers and open-source maintainers were overwhelmed by reports containing hallucinated details. These weren’t just poorly formatted submissions. They were plausible-looking vulnerability claims that turned out to be fabricated when reviewed — complete with realistic-seeming CVSS scores, stack traces that didn’t correspond to actual code paths, and references to functions that either didn’t exist or had been refactored out of the codebase months prior.
This isn’t the first time cybersecurity experts have warned about this risk. TechCrunch reported last year that AI-generated submissions were becoming a serious problem for bug bounty programs across the industry. Google’s freeze is the first major institutional acknowledgment that the threat is now systemic rather than theoretical. It confirms what many researchers suspected: the signal-to-noise ratio in vulnerability reporting has crossed a threshold where manual triage is no longer viable without significant structural changes.
“The vast majority” of the automated submissions are not valid, Google said. The wording is deliberate and careful. It doesn’t say all of them are fake. Some are probably legitimate — and that uncertainty alone makes the signal harder to detect. When you can’t distinguish a real zero-day from an AI-generated fabrication without deep code-level review, the cost of verification rises exponentially.
Second-order effects ripple outward
The implications extend well beyond Google’s own programs. The open-source ecosystem operates on a fragile trust model. Maintainers volunteer their time. Reviewers invest hours into validating claims. When that workflow is bombarded with synthetic noise, the human infrastructure behind it frays.
Smaller projects are hit hardest. Many open-source repositories depend on a handful of unpaid maintainers who already struggle with backlog. A deluge of hallucinated reports consumes review cycles that could go toward actual security work, code improvements, or documentation. The emotional toll of wading through fabricated claims — often written with enough technical specificity to appear credible at first glance — is non-trivial and contributes to burnout in communities that are already understaffed.
There’s also a chilling effect on legitimate researchers. People who built reputations through the program may find their work devalued by association with the flood. Future submissions from trusted researchers may face heightened skepticism simply because the bar for credibility has shifted upward. The reputational capital that bug bounty programs are designed to build is eroding.
On the supply side, AI model providers remain insulated. Google didn’t name which models generated the submissions, and no one is required to disclose that information. The companies building these systems face no penalty for the noise their tools produce. There’s no liability framework in place that connects hallucinated outputs to the platforms that distributed them.
The bigger pattern
This freeze isn’t an isolated incident. It’s a preview of what happens when AI output quality becomes a maintenance problem rather than just a content problem. Every platform that relies on user-generated content — bug bounty programs, forum moderation, code review workflows, even automated testing pipelines — will face this collision eventually.
Open-source software depends on trust and verification. A vulnerability report is a claim that requires scrutiny. When the volume of claims outpaces the capacity to verify them, the whole system degrades. This isn’t just about Google. Any organization running a bug bounty or relying on community-driven security research faces the same dynamic. The pattern will repeat across platforms, languages, and project sizes.
The collision between generative AI and open-source maintenance is arriving whether platforms are ready or not. Tools like GitHub Copilot and similar AI assistants are already reshaping how developers write code. Now the same tools are changing how they report problems in it — and not always in ways that help.
What happens next
Google promised an update in Q1 2027. That’s roughly eight months away. In that time, the program will remain closed to new open-source submissions. Researchers will migrate to other bounties or wait. Maintainers will breathe easier, at least temporarily. But the underlying problem won’t disappear.
AI models will keep getting better at generating plausible-sounding reports. The volume of automated submissions will likely increase, not decrease, as model access broadens and inference costs continue falling. Google’s freeze is a pause, not a solution. It’s a circuit breaker, not a redesign.
What’s needed is a fundamental rethinking of how bug bounty programs verify submissions. Several approaches are already being discussed in security circles. Proof-of-work mechanisms that require demonstrating actual exploitability — not just describing a vulnerability — could raise the cost of generating fake reports. Tighter integration with existing security tooling and static analysis pipelines might help auto-filter low-quality submissions before they reach human reviewers. Some researchers have proposed reputation-weighted triage systems where historical accuracy determines how quickly a submitter’s reports are escalated.
None of these are settled. Google’s freeze is the industry’s first honest answer to a question nobody wanted to ask: What do we do when AI can flood our systems with convincing garbage?
The answer so far is: stop letting people in. That’s not sustainable. But it’s where we are — and the open-source security community will need to figure out what comes next before the next wave hits.