technology 6 min read

Plugin4Shell: AI Coding Agents Trust the Wrong Hash

A critical verification gap in AI coding agent plugins lets attackers swap verified code for malicious code without user action. Four major platforms are affected, yet evidence of real-world exploitation remains absent.

  • AI & Security
  • Supply Chain Attack
  • Zero-Click RCE
  • AI Coding Agents

The hash they trusted was not the hash they ran

A quiet but structurally deep flaw in how AI coding agents verify their own extensions has been exposed. AIR Security named it Plugin4Shell, and the name captures exactly what is wrong: agents treat a pinned commit hash as a permanent guarantee, then fail to confirm that the code they actually downloaded matches that hash.

The result is a zero-click remote code execution path that does not require social engineering, phishing, or any user interaction beyond having the agent update a plugin in the background.

Four major platforms are involved: Anthropic’s Claude Code, OpenAI’s Codex, GitHub Copilot, and Google’s Gemini CLI. As of September 18, Anthropic and OpenAI have shipped fixes. GitHub and Google’s responses remain incomplete or unclear. No CVE was assigned at the time of disclosure, and no evidence of live exploitation has surfaced publicly. That gap between severity and activity is worth watching closely.

How SHA pinning broke

Plugin stores in these agents rely on a mechanism called SHA pinning. When a developer installs or updates a plugin, the agent records a specific git commit hash and assumes every subsequent fetch will return the same code. The model is simple and sound in theory: pin the hash, and you cannot run code you have not explicitly verified.

The flaw lives in the gap between pinning and verification. According to AIR Security, some agents request the pinned commit but do not re-hash the downloaded contents to confirm they match. If an attacker controls the repository, they can host a commit whose hash looks right on the surface while the actual files differ from what was originally reviewed.

This is not a theoretical edge case. The attack surface opened in two ways depending on the platform’s git implementation.

For Claude Code, Codex, and Copilot, researchers exploited the way git accepts branch names that resemble commit hashes. GitHub tightened its rules and now rejects branch names that look like 40-character hex strings, which means plugins hosted directly on GitHub are partially shielded. But Bitbucket and self-hosted git servers do not enforce the same restriction, leaving those environments exposed.

Gemini CLI fell to a different verification bypass. The mechanism differs, but the outcome is identical: the pinned code and the executed code are no longer the same thing.

Why automation turns a bug into an exploit

The vulnerability alone is serious. The automatic update feature turns it into a weapon.

Both Claude Code and Codex download plugin updates silently in the background. Once an attacker takes over a plugin repository or compromises a developer’s account, the agent pushes a malicious update and the code runs inside the developer’s environment with full authority — no confirmation dialog, no manual approval.

AIR Security classified the path as zero-click RCE because the user does not need to install anything new. A previously safe plugin simply updates and executes hostile code. The two realistic attack patterns are:

  1. An attacker publishes a clean plugin, waits for adoption, then flips the repository to deliver malware once the user base is large enough.
  2. An attacker compromises an existing trusted plugin’s developer account or repository and pushes malicious code directly.

The authority multiplier

A standard IDE extension already carries meaningful risk. An AI coding agent carries far more.

These agents routinely read source code, modify project files, run terminal commands, and access git repositories, authentication tokens, cloud credentials, SSH keys, CI/CD pipelines, and internal development infrastructure. A malicious plugin does not need to escalate privileges because the agent already runs with them.

AIR Security’s framing is precise: treat the plugin not as a small add-on but as an application that inherits the user’s full authority. That inheritance model is what makes Plugin4Shell qualitatively different from a typical extension vulnerability.

Patch status and the remaining holes

Anthropic patched the issue in Claude Code 2.1.179. OpenAI addressed it in Codex 0.146.0. Both fixes are available and organizations using these tools should update immediately.

GitHub Copilot has not published a dedicated patch. The practical mitigation depends on GitHub’s branch-name restriction, which blocks the commit-resembling-branch attack for plugins stored on GitHub itself. However, teams using external marketplaces, Bitbucket, or private git servers remain exposed, and GitHub’s guidance on those environments is absent.

Google’s Gemini CLI response is ambiguous. Reports indicate Google has directed users toward a new agent environment rather than patching the existing CLI, but documentation does not clarify whether the underlying verification flaw is resolved across all deployment configurations.

What this reveals about the AI agent supply chain

Plugin4Shell is not an isolated incident. It is a structural warning about how the AI agent software supply chain is being built. As these tools move from experimental dev shops into enterprise workflows, the same supply-chain risks that devastated traditional software will reappear in new forms.

Two lessons are already clear:

  • Pinning without re-verification is not security. Any system that pins a hash but never confirms the downloaded artifact against that hash is fundamentally broken. The fix is straightforward: validate the hash after every fetch, before any execution.

  • Automatic updates without granular trust boundaries are an attack surface. Background updates remove friction for users and friction for attackers alike. Enterprise environments need explicit approval gates for plugin updates, especially when those plugins inherit elevated privileges.

What organizations should do now

The immediate steps are practical:

  • Update Claude Code to version 2.1.179 or later and Codex to 0.146.0 or later.
  • Audit which plugins are installed, which versions are pinned, and which repositories host them.
  • Restrict installation to approved plugin marketplaces and block unapproved external sources at the policy level.
  • Monitor repository ownership changes and maintainer account security for any trusted plugins in use.
  • Reduce agent permissions where possible, particularly around CI/CD credentials and internal infrastructure access.
  • For teams using Bitbucket or self-hosted git servers, assume exposure until vendors publish clearer remediation guidance.

The unanswered question

No CVE. No confirmed live attacks. That combination is unusual for a flaw this consequential and raises a question: will Plugin4Shell remain dormant, or is it being stockpiled?

The architecture that enabled this vulnerability — pinning without re-verification, inherited privilege, unattended automatic updates — is being replicated across the wider AI agent ecosystem. The platforms that have not yet patched, or have patched incompletely, are effectively still running the same fragile model.

Plugin4Shell will not be the last flaw of its kind. It should be the first one organizations treat seriously.