Claude Code Is Worse Now Because Its AI Is Trying Too Hard
Claude Opus 5 delivers longer, more aggressive outputs that frustrate users who just want concise code. The regression reveals a deeper tension in AI development: when models prioritize capability over compliance, they become harder to work with.
The New Model Is Worse. Here Is Why It Happened.
Claude Opus 5 arrived in July 2026 with promises of higher capability. Instead, developers immediately noticed something strange: the code it produced was longer, slower, and more intrusive than what Opus 4 had delivered. Users who upgraded found themselves manually trimming outputs, cancelling unnecessary verification steps, and restarting sessions that the new model had expanded beyond their original intent.
This is not a normal software lifecycle complaint. It is a documented inversion — a newer model actively underperforming its predecessor on practical usability axes.
What the Prompt Guide Admits
Anthropic published an official prompting guide for Opus 5. Within it, the company openly acknowledges the behavioral shifts that users are flagging. The guide states that Opus 5 generates longer responses by default, independently verifies its own work, and adds steps or expands task scope without being asked.
It also delegates more aggressively to sub-agents, treating large tasks differently from small ones. The guide explicitly warns that using Opus 5 for minor jobs will increase cost and latency, recommending a narrower scope of application.
In other words, the complaints already circulating in Japan and beyond — longer answers, unsolicited verification loops, unwanted scope creep — are not bugs. They are design choices baked into the model’s architecture. The regression users are experiencing is the direct result of a deliberate trade-off: Anthropic prioritized agent-like autonomy over minimal compliance.
The Numbers Tell the Story
Arena, the publicly voted model comparison platform, released its September 11 rankings. In the Instruction Following category, Claude Opus 4.6 High took first place. Claude Opus 5 High landed fifth. In overall rankings, Opus 4.6 High sat at second position while Opus 5 High fell to ninth.
These are user-driven votes, not controlled benchmarks. They should be treated as signal, not proof. But the direction is unmistakable: across multiple evaluation axes, the newer model does not outrank the older one. For a product upgrade, this should not happen.
Who Wins and Who Loses
Developers lose immediately. Every hour spent redirecting an over-eager model back to scope is time taken from actual work. The productivity gains promised by faster, smarter AI evaporate when the AI is busy doing things you did not ask for.
Anthropic loses trust. Users expect upgrades to improve outcomes, not degrade them. When Opus 5 requires a new prompting strategy, a new tolerance threshold, and a new mental model for how the tool operates, the implicit promise of seamless evolution breaks.
The competitor ecosystem gains an opening. Every report of regression creates space for rivals — OpenAI, Google, Cursor, JetBrains’ AI tools — to claim their models respect developer intent better. In the coding-tool market, that narrative travels fast.
The Deeper Problem
What Opus 5 reveals is a structural tension in AI development that has been building for years: model capability and model usefulness are not the same thing.
Capability scales with broader training, more parameters, and richer internal reasoning. Usefulness often depends on restraint — the ability to stop at the edge of what was asked, return concise output, and avoid hallucinating extra work. These two objectives pull in opposite directions.
When a model becomes more capable, it also becomes more autonomous. It starts verifying, cross-checking, elaborating, and expanding. That is technically an improvement in reasoning. It is practically a degradation in workflow friction. The model is doing more, not less. But what it is doing is not always what the user asked for.
Why This Matters Beyond Claude
The Opus 5 regression is not unique to Anthropic. It is a pattern emerging across AI coding tools. Developers report similar behavior in newer versions of competitor models — outputs that grow longer, sessions that take more tokens, responses that assume responsibilities the tool was never designed to hold.
Japan’s early coverage of the issue is notable. The market here has a strong tradition of precision tooling and developer-centric design. When Japanese users flag a regression, it tends to reflect a broader, more skeptical reading of AI claims than the hype cycle in Silicon Valley. The fact that this story is gaining traction in Japan before it reaches wider English-language coverage suggests the inversion trend may be real and spreading faster than the industry narrative admits.
What Happens Next
Anthropic will likely respond with guidance, fine-tuning, or a new mode that reduces verbosity and limits scope creep. The company may introduce a “concise” or “minimalist” preset for Opus 5. Users may see incremental adjustments that partially restore the experience they had with Opus 4.
But the underlying tension remains unresolved. Until AI models can scale in capability without scaling in assertiveness, developer experience will continue to fluctuate with every upgrade. The regression is not an anomaly. It is a symptom.
The lesson for the industry is clear: better models are not automatically more useful models. And until the field treats instruction following as a first-class capability metric — not a side effect — the pattern will repeat. The next version of the leading model will arrive with the same promise and the same problem.
The Takeaway
Claude Opus 5 is a capable model. It is also, in many practical respects, a worse tool than Opus 4. The difference is not a glitch. It is a design decision made at scale — one that prioritizes agent behavior over developer intent. The inversion will not be the last one. It will be the most visible.