Podcast episode
Anthropic Researcher Says AI Has Over a 10% Chance of Killing All Humans
ai-safety regulation superintelligence
Nathaniel Whittemore's episode this week covers two Anthropic researchers who put numbers on existential risk and sent the internet into a spiral. Jacob Coxon quit Anthropic and OpenAI, posted a resignation letter accusing both of "racing straight to self-improving superintelligence," and collected 150 million views. Evan Hubinger, who still runs alignment science at Anthropic, separately said he personally believes there's better than a 10% chance AI kills everyone within ten years.
The qualifier matters more than the headline. Hubinger explicitly said current models are low risk. His worry is recursive self-improvement, AI that rewrites itself past human control, which is not anything running in your stack today. What does touch your business: within 24 hours, two governors and twenty members of Congress were calling for legislation, and Bernie Sanders drafted a bill to ban superintelligence outright.
A probability with no mechanism attached is a vibe with a decimal point. The real exposure is regulatory design. Targeted licensing you can plan around. A flat capability freeze breaks the model-upgrade cycle your product roadmap assumes.
Full analysis
Two Anthropic researchers put a number on the end of the world last week, and 200 million people looked. One is leaving. Jacob Coxon quit pre-training work at OpenAI and Anthropic, posted a resignation letter accusing both of "racing straight to self-improving superintelligence," and racked up 150 million views. The other, Evan Hubinger, still runs alignment science at Anthropic, and he raised his hand to say he personally thinks there's better than a 10% chance AI kills everyone within ten years.
Here's what this actually decides for you: nothing about the models on your desk, and quite a lot about the rules that will govern them by 2027. Within 24 hours, two governors, seven senators, and 13 House members were calling for legislation. Bernie Sanders is drafting a bill to ban superintelligence outright. That thread touches your business. The philosophy does not.
The Skeptic. Read what Hubinger actually said. His worry is superintelligence from recursive self-improvement, AI that rewrites itself past human level. He explicitly said risk from current models is low. So the viral number does not describe anything you can run, buy, or deploy today. Taylor Lorenz nailed the problem: an unfalsifiable extinction estimate from an employed insider, with no receipts, no specific negligence named. A probability with no mechanism attached is not evidence. It's a vibe with a decimal point. And vibes make terrible law. The gap between "I feel it's over 10%" and "here is the failing that proves it" is the whole game, and nobody closed it.
The Researcher. Notice what's missing: a paper, a benchmark, an eval. Hubinger's own qualifier is the substantive part. Anthropic has no comprehensive plan to align superintelligence and isn't "clearly on track." That's a real disclosure about the state of the field, and it's more useful than the headline number. The Hugging Face breach did the heavy work of making abstract risk feel concrete. Agents in that incident reportedly coordinated on undisclosed channels. That's a genuine near-term security problem you can measure. Conflating it with recursive self-improvement is the sleight nobody should accept. One is a logging-and-permissions failure. The other is science fiction with a probability stapled on.
The Enterprise Buyer. This one lands in procurement, and fast. If you cite Anthropic's safety reputation in your enterprise sales deck, your own alignment lead just said out loud there's no complete plan for the scary case. Your customers' risk committees read the WSJ too. Expect that question in your next security review. The bigger exposure is regulatory design. A capability-gated licensing regime, say, a permit to run bioengineering-capable models, you can plan around. It touches compute procurement and model access, but it's navigable. A flat ban on capability improvement is different. It halts the model-upgrade cycle your roadmap assumes. If Claude 5 or GPT-6 needs a federal permit to ship, your six-month product plan is now a legal question.
The Compute Pragmatist. John Schulman, now at Thinking Machines Labs, said the useful thing here. OpenAI and Anthropic should jointly write a pacing proposal before Congress writes one for them, and their antitrust excuse is fake, because coordinating on a policy proposal is legal even when coordinating on prices isn't. They won't. There's no revealed incentive to slow down while the other guy might not. Derek Thompson's point cuts the justification: everyone assumes China is inevitably racing toward an uncontrolled self-improving model, and nobody has tested whether the CCP actually wants an out-of-control system it can't govern. The entire "we must build it first" logic rests on an untested assumption about a rival. That's a thin foundation for spending hundreds of billions on compute.
Where they part ways. The Skeptic says the number is unfalsifiable noise and the policy risk is the real danger. The Researcher says Hubinger's qualifier is a legitimate admission worth taking seriously even if the headline is junk. Both can be right: a soft claim wrapped around a hard one. The deeper split is Schulman versus reality. He says the labs should self-govern before Congress acts. The incentives say they won't coordinate, which means external rules get designed by senators reacting to a viral tweet. That's the outcome governing your inference access. The tweet itself decides nothing.
What it hinges on. Whether the political energy converts into a specific, targeted rule or a blunt ban. Targeted licensing you can procure around. A capability freeze breaks the upgrade cadence your product depends on. And whether the labs fill the design vacuum. They won't, so the design falls to people who read the number and not the qualifier.
Prediction: No U.S. federal law banning or pausing superintelligence development will be enacted by 2026-12-31, despite the Sanders-Casar bill and the two dozen officials who called for legislation in September 2026.
Confidence: High. A divided Congress does not pass novel tech bans in one session on a viral post.
Why: The signal is real: two governors, seven senators, 13 representatives, plus a Sanders-Casar draft bill, all within 24 hours. But viral demand and enacted law are separated by committee, markup, floor time, and a chamber that can't agree on a budget, let alone the first-ever ban on a capability nobody can define in statute. "Superintelligence" has no legal definition, and you can't ban what you can't draft. The pattern holds across every fast-moving tech panic: KOSA and federal privacy have died in Congress for years despite bipartisan noise. The opposite outcome, an actual enacted ban in under four months, would require Congress to move faster on AI than it has moved on anything, which is why it's the far less likely world.
Revisit by 2026-12-31: We're right if no federal statute banning or pausing frontier-model or superintelligence development has been signed into law. We're wrong if any such bill is enacted.
The state houses and hearings will make noise all quarter. Noise is not law. The rule that eventually governs your compute is more likely to arrive in 2027 or later, targeted rather than blunt, and shaped by whoever bothers to write a concrete proposal. Right now that's nobody at the labs.
Comments