technology 5 min read

Google's Gemini 4 Passes the Test, Fails the Job

Internal Google reports reveal Gemini 4 Argon's coding performance lags behind its benchmark scores — a problem the AI industry calls 'benchmaxxing' that may reflect a wider failure of Western AI tools to capture real-world East Asian development workflows.

  • Artificial Intelligence
  • OpenAI
  • Anthropic
  • Google Gemini
  • AI Benchmarks

Benchmarks Are Lying to You

Google showed off Gemini 4 Argon to a small circle of cybersecurity partners last week. On paper, it looks unbeatable — top-tier scores across multiple benchmarks, reportedly beating OpenAI’s Astra on at least one security evaluation. Inside the company, the narrative is confidence. One engineering lead told Bloomberg there is a “broad consensus” that Gemini 4 is state-of-the-art.

But another cluster of Google employees, including some with direct access to the model’s development pipeline, are telling a different story. The benchmarks don’t match the experience. Coding performance, in particular, is uneven. Front-end design work — building app interfaces, webpage layouts, user experience flows — is where the gaps show most clearly. A model that scores well on standardized tests produces code that developers find frustrating to work with in practice.

The article’s Korean headline captures it perfectly: “Why does it feel frustrating to use?”

This is not a Gemini-specific problem. It is the central tension of the current AI industry, and Korean developer feedback on Gemini 4 puts a sharp regional lens on something that matters globally.

The Benchmaxxing Trap

The AI industry has a word for this now: benchmaxxing. It describes the growing incentive to optimize models for benchmark scores rather than for the messy, unstandardized tasks that real users actually face. Benchmarks are visible. They can be ranked. They drive sales pitches. The result is that model builders pour effort into squeezing points out of tests, not into building tools that feel good in a developer’s hands.

Edwin Chen, founder of AI startup Surge AI, compared it to the SAT. A student can score perfectly on a practice exam and still struggle in a college classroom. A model that dominates MMLU or HumanEval does not necessarily produce clean, production-ready code on the first try.

Gemini 4 is not universally weak. Sources familiar with internal evaluations say it handles video metadata extraction well, performs solidly on safety and cybersecurity tasks, and communicates in ways that feel natural and clear. These are real strengths. But in a market where the differentiator is becoming actual coding productivity — not theoretical capability — the gaps matter.

What Korean Developers See That Others Miss

The concern surfacing in Korean tech circles about Gemini 4 is a reminder that Western AI tools often optimize for workflows and language patterns that do not translate well to East Asian markets. Korean developers write code in English but navigate documentation, frameworks, and design conventions shaped by a very different user ecosystem. A model trained primarily on English-language repos and American dev practices may produce syntactically correct output that misses the practical nuance of how those tools are actually used in Seoul, Tokyo, or Singapore.

This is not unique to Korea. It is a systemic blind spot. When the dominant benchmarking ecosystem is American, when the largest pool of training data comes from English GitHub repositories, and when the primary customer base for coding AI is Silicon Valley, the resulting models will reflect those biases — even if the benchmark scores look neutral.

Gemini 4’s struggles with front-end work, where design conventions and user experience expectations vary significantly across regions, may be the canary. If a model optimized for Western dev patterns stumbles on interface design, it likely has similar blind spots in Chinese-language frontend ecosystems, Japanese SaaS workflows, and other non-Anglo markets that have not yet made the same noise.

The Competitors Are Not Standing Still

Inside Google, some employees are watching Anthropic’s Fable and OpenAI’s Astra closely and concluding that both are advancing faster than Gemini. That perception matters regardless of whether it is entirely accurate. In a market where developer loyalty is fragile and switching costs are low, being perceived as behind is nearly as damaging as being behind.

OpenAI and Anthropic are pushing hard into coding agents and integrated development products. Google has the distribution advantage — Gemini is already embedded in Gmail, Search, Chrome, Maps, and other services with over a billion users each. But distribution alone does not win the model race. If developers find that their favorite AI coding assistant consistently produces mediocre front-end code while the alternatives deliver better results, they will migrate. Google’s installed base is an asset, but it is not a moat against worse product experience.

The Cost of a Missed Launch

The stakes are higher than most reporters are making them out to be. Google previously announced Gemini 3.5 Pro in May and then quietly dropped it before launch. Manavik Sandhu, an analyst at Bloomberg Intelligence, estimated that training a single large model of this size can cost up to $400 million, not counting the salaries of the research staff involved. That is money gone with nothing to show for it.

Gemini 4 itself is described as a very large model with a context window of up to 1 million tokens — roughly 750,000 words. Large models are expensive to run. Every additional token of context increases inference costs, and Google’s own services — Search, Gmail, Cloud — will eventually need to absorb those costs if Gemini 4 is to power them at scale.

The question is whether the model delivers enough real productivity gains to justify the economics. Benchmarks say yes. Developers using it day to day are less convinced.

Why This Matters Beyond Google

The Gemini 4 episode is a case study in a problem that will shape the entire AI industry for years. The gap between benchmark scores and real-world performance is not a glitch — it is a structural feature of how the industry currently measures success. Until that changes, the models that look strongest on paper will keep disappointing in practice, and the companies that ignore the discrepancy will lose developer trust faster than they expect.

Korean developers are among the first to flag this because they are among the most active non-English-speaking AI coding communities. Their feedback is a signal that the bias toward Anglo-centric benchmarks is real, measurable, and costly. If Gemini 4 cannot win over Korean frontend developers, it will have a harder time winning over any developer who does not fit the profile of the typical American tech worker.

Google’s internal debate is not yet resolved. But the fact that it is happening inside one of the world’s largest AI labs should be a warning to everyone betting on benchmark scores as proof of product readiness.