ChatGPT Co-Founder Says the LLM Era Is Over. He Built the Alternative.
Diogo Almeida, who co-invented ChatGPT, says the token-by-token paradigm is hitting its limits. His new model Jev claims GPT-5.6-level performance at a fraction of the cost by abandoning autoregressive generation entirely. What's actually different—and why it matters for everyone building on AI.
The question that drove a ChatGPT co-founder out the door
Diogo Almeida spent years at OpenAI helping build ChatGPT. When he left, he didn’t start another startup chasing better prompts. He built something deliberately different.
Almeida’s frustration was simple and, on reflection, under-discussed: superhuman chat models did not lead to AGI. Or anything close to it. The architecture itself—the next-token-prediction engine that became the default for every major AI lab—appears to be a bottleneck, not a bridge.
Two years of stealth work later, he’s releasing Jev, the first “System One Model” from his company TypeSafe AI. It’s not an LLM. That distinction matters more than the press release suggests.
What Jev actually does differently
Standard LLMs generate output one token at a time, left to right, in a sequential chain. Each token depends on every token before it. That’s why they can write novels, code, and structured documents—they’re essentially continuing text. It’s also why they’re slow and expensive to run. Computation compounds with every token generated.
Jev was built for the opposite purpose: it doesn’t generate text. It takes a task specification and outputs a pre-defined structure in a single parallel pass. Think of it less like a writer and more like a compiler.
Almeida himself noted the historical parallel: replacing sequential computation with parallel computation is the same structural leap that made Transformers overtake RNNs. The industry is doing that again, but inside the transformer paradigm, rather than stepping outside it.
The numbers are aggressive
Here’s what TypeSafe AI published:
Input cost: $0.042 per million tokens. Output: $0 — described as “too low to measure.”
On benchmarks matched against GPT-5.6 Terra-class models, Jev claims equivalent quality at roughly 1/100th the cost. Tool-call error rate: 0%. Every output includes a confidence score.
The demonstration included a Wikipedia navigation race—starting from one page and reaching a target page using only hyperlinks. Jev finished faster than all competing models, even those running in optimized non-inference mode. Almeida attributed the margin partly to the absence of hallucination compounding: when you’re choosing between hundreds or thousands of links, errors multiply with every hop.
Perhaps the most striking demo: Jev played DOOM for one hour at a cost of $7. For comparison, running a comparable LLM-based agent for an hour of continuous decision-making typically costs far more, and with no guarantee of coherent action.
Who wins, who loses
The immediate winner is anyone building AI agents that perform structured tasks—code generation, data extraction, tool orchestration, game-playing agents, automated workflows. These are exactly the applications where the token-cost math of LLMs has been the biggest friction point. At $0.042 per million input tokens with near-zero output cost, the economics flip.
The loser is any product or workflow that was priced around LLM inference costs. If your business model assumed users would pay per token for autonomous agent tasks, Jev undercuts the entire cost structure by two orders of magnitude. That’s not competition. That’s a ground collapse.
But there’s a catch worth taking seriously.
The limitation is the trade-off
Jev cannot generate open-ended text. It produces pre-defined structures. That means it’s not a general-purpose replacement for ChatGPT-style interaction. It won’t write a story, draft an email, or engage in free-form reasoning.
This is not a flaw in the description—it’s the design constraint. The same parallel architecture that delivers speed and cost advantages means Jev operates in a narrower bandwidth of capability. It’s purpose-built for tasks where the output format is known in advance and correctness is measurable.
That’s also why Almeida’s framing is careful: he positions Jev as complementary to LLMs, not a full replacement. The future may not be one or the other but a layered stack—LLMs for open-ended reasoning and creative tasks, System One models for execution.
What this signals
The fact that a ChatGPT co-founder is publicly betting against the autoregressive paradigm is noteworthy regardless of Jev’s final performance. It signals that the people who built the current architecture see its limits more clearly than anyone else—and they’re not shy about acting on that insight.
The RLCD training method Almeida references (reinforcement learning via constrained decoding, reportedly) is another data point. It suggests the next wave of advances may come not from scaling existing architectures but from fundamentally different training paradigms.
Korean and Japanese tech readers have been tracking this angle for some time. The region’s enterprise AI adoption—particularly in automation and robotics—has always been more comfortable with structured, deterministic systems than with probabilistic text generators. Jev fits that instinct.
What happens next
Three things to watch:
First, independent benchmarking. TypeSafe AI’s numbers are impressive but self-reported. The industry will need third-party validation before pricing models and investment theses shift.
Second, integration patterns. If Jev handles structured tasks at scale and LLMs handle open-ended reasoning, the real question is how they’re wired together. The middleware layer between them becomes the new platform opportunity.
Third, the cost signal. Even if Jev’s current capabilities are narrow, the economic pressure it creates is real. Every developer watching their inference bill will ask: why am I using an LLM for a task that a System One model could do for 1/100th the price? That question spreads fast.
Almeida left OpenAI because he believed the current path doesn’t lead where we need it to go. Whether he’s right will depend on what Jev can do once the demos end and real workloads begin.