business 6 min read

Why Anthropic Defectors Are Changing the AI Safety Debate

Two Anthropic researchers gone in weeks, both sounding the same alarm. The insider exodus is shifting AI safety from fringe concern to boardroom crisis.

  • Anthropic
  • AI Regulation
  • AI Safety
  • AI Alignment
  • Existential Risk

The Warning That Won’t Shut Up

Evan Hubinger posted on X that there is more than a 10 percent chance AI could extinguish humanity within a decade. Two days later, Jacob Colson — fresh off his departure from Anthropic — described a near-term trajectory toward superhuman systems that could hack anything and reshape every industry overnight. Both men worked on AI alignment at the company. Both are gone now, or at least changing their job titles.

The coincidence matters. This is not a lone dissenter blowing the whistle from the outside. It is a pattern forming inside the lab itself.

From Fringe to Foreground

The AI safety conversation has undergone a quiet but real shift over the past year. Five years ago, worrying about existential risk made you a cautionary figure on the internet, someone to be politely sidelined. Today, Hubinger and Colson hold positions that were literally built to take that worry seriously. They are not hobbyists. They are the people Anthropic hired to make sure its models do not accidentally unmake the world.

And yet they are leaving, and they are speaking publicly, and they are saying the same thing: the alignment problem is not solved, the pace of development is outrunning the safeguards, and the companies are not going to stop themselves.

The debate has moved from whether AI poses an existential risk to how large that risk actually is. That may sound like a technical adjustment. It is not. It is a generational change in how seriously the industry takes its own creations. The 10 percent figure is not a precise probability estimate. It is a signal flare. Someone who knows the architecture is looking at the trajectory and deciding the slope is steep enough to worry about.

The PR Question, and Why It Misses the Point

Dame Wendy Hall, the UK government’s AI advisor, suggested on BBC Radio 4 that the warnings might carry a promotional element, given that both Anthropic and OpenAI are expected to go public soon. It is a reasonable line of questioning. Corporate actors routinely use fear to sharpen their positioning.

But dismissing insider warnings as marketing misses something important: these researchers are burning social capital, not building it. When you leave a well-funded, high-status lab and go on X to say your former employer is moving too fast and not going far enough on safety, you are not acquiring fans in Silicon Valley. You are making enemies. The fact that they are doing it anyway matters more than any press cycle it might generate.

The AISI Story

A more concrete episode came out of the UK. The Financial Times reported that Anthropic did not submit its latest model to the AI Security Institute for evaluation. Anthropic declined to comment. The UK Cabinet Office said it continues to work with industry partners including Anthropic, which is exactly the kind of sentence you write when you want to sound cooperative while saying nothing at all.

This is the new normal of AI safety governance: voluntary submission, corporate silence, and a government that prefers the appearance of partnership over the friction of enforcement. The UK has an AI Safety Institute. It exists. It has no teeth.

What Anthropic Actually Says — and What It Withholds

Anthropic’s own August safety report admitted something that should disturb anyone paying attention: confidence in the company’s own safety assessments is declining. The report stated that risks of model misuse and of advanced AI carrying out automated research leading to catastrophic outcomes were still judged low, but flagged “early signs of acceleration” that could change the calculus. The phrasing is careful. The meaning is not.

The company also stopped short of providing its most capable model to the UK’s AISI. That omission is a choice. It is not illegal. It is simply the kind of choice that looks very different depending on whether you are an investor or a researcher who thinks the world might end sooner than expected.

The Defections

Colson’s trajectory is notable. He worked at OpenAI before moving to Anthropic. He has now left both. His post announcing the departure described both companies as failing to act responsibly, predicting that superhuman AI systems would emerge within reach. Hubinger’s post, seen over ten million times, went further in numerical specificity: a 10 percent risk of human extinction within ten years, no clear plan for solving alignment at superintelligence scale, and a description of the current trajectory as unclear rather than aligned.

These are not abstract worries. They are coming from people whose entire professional mandate was to prevent exactly this outcome.

The Growing List

Hubinger and Colson are not isolated. OpenAI’s chief scientist, Yakub Pachocki, called for an extremely cautious approach this month and warned that continued human control may require intervention beyond what the field is currently willing to do. Anthropic’s own CEO, Dario Amodei, and fellow executive Jared Kaplan have publicly pushed for slowing development. A letter signed by over 1,300 AI industry workers urged the US government to help build technical and governance tools to deliberately pace automated AI development.

The message is consistent across organizations, across roles, across the ideological spectrum of the field: the speed is the problem, and the people closest to the work are the ones seeing it most clearly.

What Changes Next

The immediate consequence of these warnings is political. Darren Jones, a senior UK official, has written to the prime minister calling for a new multilateral treaty on AI safety. The logic is straightforward: if the developers inside the companies will not self-restrain, and the companies themselves are racing toward IPOs and market dominance, then governments need to create binding constraints before the technology creates facts on the ground that no treaty can undo.

The harder question is whether any treaty can actually constrain a technology that can be developed in parallel across multiple jurisdictions by private actors with no shared incentive to slow down. Treaties work when verification is possible and when defection carries a cost. Neither condition is easy to satisfy with frontier AI models, which live in code, not in silos.

The Real Shift

What is happening inside Anthropic and its orbit is not a scandal. It is an institutional diagnosis. The people hired to keep the technology safe are telling the public that the technology is moving faster than the safety work. That is not a critique of one company. It is a critique of an industry structure where the primary competitive pressure is to ship first, not to ship safely.

The 10 percent number will be debated. It should be. Probability estimates from insiders are not scientific findings. But the number is not the story. The story is that the insiders are leaving, they are speaking in public, and they are saying the same thing from different exits. That pattern is harder to dismiss than any single data point.

Investors will decide whether to listen. Governments will decide whether to act. The researchers are past the point of deciding for anyone else. They have already made their choice.