science 5 min read

Sutton's 20-Watt Bet: Why a Trillion-Parameter Brain Could Upend AI

Richard Sutton says today's AI is stuck — learning stops the moment models deploy, and the industry's answer to data hunger is just a big mistake. His counter: continual learning at brain-like power levels. The question is whether Oak Lab can prove it.

  • Richard Sutton
  • Continual Learning
  • AI Efficiency
  • LLM Criticism
  • Oak Lab
  • AI Paradigm

The 20-Watt Bombshell

Richard Sutton just told the AI world that everything it’s been building for the past four years is a detour.

The citation heavyweight behind Reinforcement Learning — the man who gave DeepMind its theoretical spine and penned the essay many call the field’s intellectual manifesto, “The Bitter Lesson” — has declared the current LLM-centric paradigm fundamentally broken. And he’s not just diagnosing. He’s left. Through a Sequoia Capital podcast, Sutton announced the founding of Oak Lab, a startup built around what he describes as a radically simpler idea: intelligence doesn’t need megawatts. It needs continual learning at 20 watts.

That number — 20 watts — is the power budget of a human brain. Your refrigerator draws more. A data center housing a single large language model draws enough to power a small town. Sutton’s argument is blunt: if you can do this with 20 watts, the entire compute arms race is unnecessary.

The Diagnosis: Learning That Stops

Sutton’s sharpest jab targets something most users never notice — because it’s the default state of every deployed LLM today. Once a model is frozen and shipped, it learns nothing. It doesn’t adapt. It doesn’t improve. It doesn’t forget, either, which is another way of saying it doesn’t live.

This is what Sutton calls the “fatal flaw”: systems claiming PhD-level expertise that cease accumulating knowledge the moment they enter production. They are, in effect, talking monuments.

His alternative is “continual learning” — models that continuously update their weights through interaction with their environment, the way a brain does. Not fine-tuning on a static dataset. Not RLHF with human preferences. Actual, ongoing weight updates driven by environmental feedback.

The obvious objection is catastrophic forgetting — the well-known problem where a neural network, upon learning new information, overwrites previously acquired knowledge. Sutton acknowledges this directly. His proposed solution is an algorithm his team calls “continual backpropagation,” which assigns individualized learning rates to individual network weights, preserving existing knowledge while injecting stochastic new units that enable ongoing adaptation.

It sounds plausible in theory. The question is whether it scales.

Synthetic Data as a Trap

Sutton’s critique extends to the industry’s current workaround for data exhaustion: synthetic data. He dismissed it without ceremony — “just a big mistake” — and his collaborator Kuram Zaveed backed the reasoning. Validating synthetic data eventually requires human expert judgment, which creates a bottleneck that collapses the whole economic model.

This is where Sutton’s “Big World Hypothesis” does the heavy lifting. The physical world is infinitely more complex than any simulation a human could construct. Human-made simulations, by definition, are bounded by human understanding. You cannot synthesize your way out of a complexity problem whose source you cannot fully model.

The implication is uncomfortable for every AI company currently betting billions on data generation pipelines. If Sutton is right, those pipelines don’t solve the data problem — they just move it one step downstream onto human labor that will itself become the constraint.

Who This Upends

The Sutton-Zaveed position strikes at the center of the dominant AI narrative: that scale is the only path to intelligence. Every major lab is currently racing to build bigger, train longer, and spend more. Their justification is that intelligence is a function of compute — more parameters, more tokens, more FLOPs.

Sutton says that’s a category error. Intelligence, as he’s argued since “The Bitter Lesson” in 2019, is a function of algorithms that exploit computation — not of computation itself. The bitter lesson was always that methods which automate search and learning through general-purpose algorithms consistently outperform hand-engineered approaches. The current paradigm, he’s now saying, is the hand-engineered approach on a massive scale.

For Big Tech, the threat is existential in a way most press coverage misses. If continual learning can achieve what massive pre-training claims to achieve — and do so at a fraction of the energy cost — then the moat built on GPU capacity and data access evaporates. The advantage shifts from those who can spend the most to those who can learn the best.

That’s why the Sequoia connection matters. This isn’t an academic exercise. Oak Lab is being built as a company, not a research group.

The Hard Part: Proving It Works

Zaveed is honest about the engineering gap. Storing a trillion parameters at 20 watts is currently impossible given memory architecture constraints. The prediction is 5 to 10 years. That’s a long runway in AI time, where a year can look like a decade.

Sutton also made a provocative framing choice during the discussion: when critics call his stance “radical,” he replied that he’s holding to the plain, standard view and the field has gone insane. He used the example of a squirrel navigating physical obstacles to illustrate that sensorimotor intelligence — the kind requiring real-time interaction with a complex world — remains far beyond current language-based systems. Language, he argues, accounts for only about 20 percent of true intelligence.

This is a direct challenge to the LLM-as-AGI narrative that dominates the industry. It places Sutton firmly in the camp that views embodied, interactive learning as the missing ingredient — a position shared by a minority of researchers but largely ignored by the capital markets.

What Happens Next

The immediate impact is limited. Oak Lab is early. The algorithmic claims haven’t been peer-reviewed or independently reproduced. The 20-watt target is aspirational. No one should read this as a near-term threat to the compute-heavy paradigm.

But the longer-term signal is clear. Sutton is one of the most respected theorists in reinforcement learning and has never been afraid to challenge orthodoxy. His participation in a startup founded on an alternative architecture gives that alternative more credibility than most new approaches receive. If Oak Lab demonstrates even partial success — a model that learns continuously without catastrophic forgetting at a meaningful efficiency gain — the entire investment thesis for massive-scale training faces scrutiny.

The 20-watt target is almost certainly optimistic for the timeline offered. But the direction is unambiguous: the question isn’t whether the field will eventually incorporate continual learning principles. It’s whether the current paradigm will be dismantled or merely augmented by them.

Sutton would say we already know the answer to that. His critics are still writing their papers.