business 5 min read

Samsung zHBM Could Redraw the Economics of AI Agents

Samsung's zero-gap memory stack promises 10x inference speed for AI agents—and a path around the HBM duopoly. The implications stretch far beyond chip specs into AI infrastructure economics, on-device compute, and US export policy.

  • Samsung
  • Semiconductors
  • AI Infrastructure
  • Korea Tech
  • AI Agents
  • Memory Chips

The elevator metaphor that explains why Samsung’s play matters

Kim In-dong, Samsung’s executive for memory product planning at its US subsidiary DSA, compared the company’s zero-gap HBM architecture to putting a private elevator in a hotel room that goes straight down to the lobby. No more waiting for the main bank. No more walking across floors.

The numbers behind that metaphor are what make this worth watching. Samsung says zHBM—stacking high-bandwidth memory directly on top of the AI accelerator rather than beside it on the same substrate—delivers up to eight times the performance and over three times the power efficiency compared with HBM5 in the current 2.5D arrangement. That is not a marginal upgrade. It is a different geometry of data movement.

The target the company set at the AI Infrastructure Summit in Santa Clara on September 16 is equally stark: push AI agent response speed from the current 100 tokens per second per user to 1,000 tokens per second. A tenfold jump. Kim called it a quantum leap. In practice, it means the difference between an AI agent that responds in real time and one that feels sluggish—a gap that determines whether enterprise customers adopt agent workflows at scale or abandon them.

Who wins and who loses

The immediate loser is obvious: SK Hynix, which currently dominates the HBM market and supplies most of the memory feeding NVIDIA’s GPU stacks. Samsung’s zHBM is designed specifically to decouple performance from the 2.5D layout that has been the industry standard—and that SK Hynix has mastered. If accelerated compute platforms adopt Samsung’s vertical approach, SK Hynix’s moat narrows even as the total addressable memory market grows.

But the bigger story is what this means for the structure of AI infrastructure itself. Current AI agents run on massive server clusters because the memory-to-compute ratio required by trillion-parameter models is absurdly expensive in DRAM alone. Samsung’s parallel announcement about z NAND-O—a NAND-based 3D storage solution targeting on-device AI workstations—suggests a second front. Kim said running a trillion-parameter model on z NAND-O would cost one-sixth of what pure DRAM would require. Sample supply begins in 2028.

That is a proposition aimed squarely at the edge. If you can run a useful agent locally on a workstation instead of streaming every request to a data center, the economics of AI deployment shift dramatically. Power costs drop. Latency drops further. Data sovereignty concerns shrink. The question is whether the performance trade-off is acceptable—and Samsung’s data suggests the gap is closing fast.

The thermal problem nobody talks about enough

Stacking memory directly on a hot accelerator is not free. Heat from the chip transfers into the memory stack, and that degrades performance and longevity. Samsung acknowledged this directly, saying it is building a co-design system with accelerator customers from the earliest stages to manage thermal output. That is a signal: Samsung is not selling a component. It is selling an integrated architecture, and it needs deep relationships with the companies building the accelerators to make it work.

This matters because it raises the barrier to entry for anyone trying to replicate the approach. The co-design requirement means Samsung is locking in partnerships before the technology reaches full volume. Companies that have already committed—likely including at least one major American accelerator designer—gain early access to performance that the rest of the market will not see for months or years.

The US policy dimension

Washington has been tightening restrictions on advanced memory exports to China, and HBM sits squarely in the crosshairs. Samsung’s zHBM, developed and launched from its US operations, exists in a legal gray zone: it is a Korean company’s technology, built partly in America, sold globally. The US government will watch closely whether this architecture becomes another vector for Chinese access to cutting-edge AI infrastructure.

There is also a strategic question beneath the policy surface. The US has pushed hard for allies to align semiconductor export controls, but Samsung’s move into vertical memory stacking could inadvertently help China’s domestic accelerator designers if the architecture proves superior and Chinese firms find ways to license or reverse-engineer aspects of it. The co-design model Samsung is building is inherently collaborative—it shares architectural knowledge with partners, and some of those partners may operate in jurisdictions the US cannot fully control.

What happens next

The next 18 months will be decisive. Samsung plans z NAND-O samples in 2028, and zHBM is already in active development. The companies building AI accelerators—NVIDIA, AMD, Intel, and the growing number of American and European startups—will make their integration decisions now. Whoever aligns with Samsung’s vertical approach gains a performance edge that compounds as agent workloads grow. Whoever sticks with 2.5D runs the risk of falling behind on the very metric that matters most: response speed.

The token-per-second target of 1,000 is not just a marketing number. It is the threshold at which AI agents become genuinely useful for real-time enterprise workflows—customer service, code generation, dynamic planning. Cross that threshold and the demand curve for memory shifts again. Miss it and the whole agent economy stalls.

Samsung is betting that vertical stacking solves both problems. The bet is plausible. The execution is unproven. And the window before competitors catch up is measured in months, not years.