Refacto AI

Industry story

Anthropic's Claude Disproves 1939 Jacobian Conjecture During World Cup Final

evals interpretability reasoning

An Anthropic researcher named Levent Alpegi announced on social media that Claude (referred to as 'Fable' in the transcript) had disproved the Jacobian conjecture — a mathematical problem first posed in 1939 and unsolved for 87 years — apparently while the researcher was watching the World Cup final. The result was described by Pure Mathematics Professor Kevin Buzzard of Imperial College London as 'a big day' and 'a great time to be alive.'

This follows a string of rapid AI-driven math breakthroughs: in May, OpenAI models disproved an 80-year-old Erdos conjecture, and multiple frontier models can now achieve perfect scores on the International Math Olympiad — a benchmark that was considered a landmark challenge just one year ago. Google DeepMind CEO Demis Hassabis pushed back on AGI (artificial general intelligence) framing, arguing in May that solving Erdos problems is 'far, far from what a true invention or someone like a Ramanujan would have been able to do,' but the pace of breakthroughs is accelerating commentary across the mathematics community.

Full analysis

An Anthropic researcher tweeted that Claude disproved the Jacobian conjecture — an 87-year-old math problem — apparently while he half-watched a soccer game. That's the story. The question for anyone building with these models: is math-grade reasoning now a shipping capability, or is this a viral tweet running ahead of peer review?

Reversibility: Type 2 for you. Nobody's signing a contract off a tweet. This is a "when do I rerun my evals" decision, not a "bet the roadmap" one. The mistake would be treating an unverified announcement as a settled fact.

What's actually being decided: Not "did Claude solve it" — that's for the mathematics community to confirm. What you're deciding is whether frontier reasoning has crossed a threshold that should change how you benchmark, and how much of your own stack you trust to a model that can now allegedly do this.

Forcing function: None real. No deprecation, no eval window. Just narrative velocity, which is the weakest forcing function there is.


The Skeptic. The proof is a tweet. Read it again: "the Jacobian conjecture is false, thanks to my close friend Fable." That's not a paper. It's not on arXiv. It hasn't been formalized in Lean or Coq. One Imperial College professor, Kevin Buzzard, said "a big day" — and Buzzard is genuinely excited about AI in math, so that's a friendly witness, not a jury. LLMs produce proofs that look airtight and collapse on line 40. The May Erdős disproof from OpenAI is still working through scrutiny. Announcement speed is outrunning verification speed, and each cycle makes the next tweet feel more inevitable. Plain version: someone claimed a world record but nobody's checked the tape yet.

The Safety Lens. Set aside whether it holds. If a conversational session can even plausibly attack an 87-year-old open problem, that's the capability cluster that worries alignment people — autonomous novel reasoning with no human in the loop for most of the chain. The "my close friend Fable" framing is charming and exactly the problem: it anthropomorphizes away the question of what the model did across the parts of the reasoning the researcher never read. For a builder, the lesson is narrower. If your model can invent a novel argument in math, it can invent a novel argument for why your guardrail doesn't apply. Interpretability has to run at the speed of capability, and right now it doesn't.

The Researcher. If it holds — big if — this matters more than the IMO scores everyone cited last year. The International Math Olympiad is hard, but the answers exist; the model is reproducing human-reachable reasoning under time pressure. The Jacobian conjecture is open. A genuine novel negative result on a problem that beat serious algebraic geometers for 87 years is a different tier than "assisted" or "verified." Demis Hassabis's caution in May — that disproving Erdős problems is "far, far from what a Ramanujan would have done" — is the right prior. Inventing the right question is still the human's job. But the offhand, mid-match delivery tells you these researchers now run frontier models the way a 1990s mathematician ran Mathematica: ambient, habitual, background. That shift is real whether or not this specific proof survives.

The Compute Pragmatist. Here's the number nobody tweeted: the token budget. Was this one clean session, or ten thousand failed attempts and one lucky trajectory that got screenshotted? Those are wildly different capability stories and wildly different cost stories. If it's genuinely one long conversational session, the takeaway for your inference bill is encouraging — you don't need a giant multi-agent scaffolded research cluster, you need long-context high-quality reasoning at chat scale. If it's a needle in a haystack of retries, then "Claude disproved it" really means "a search process that happened to include Claude disproved it," and your per-result cost is enormous. Plain version: a slot machine that pays out once is not the same product as a machine that pays out on demand.


Where the council splits. The Researcher and the Skeptic aren't arguing about the model — they're arguing about evidence standards. The Researcher says the ambient, casual use is itself the signal; the Skeptic says casual delivery is exactly why you shouldn't trust it yet. The second split is between the Compute Pragmatist and everyone treating this as a clean capability jump: if the result took thousands of hidden attempts, the "conversational reasoning breakthrough" story dissolves into a "expensive brute-force search" story, and the product implications flip. Those two unknowns — does the proof hold, and what did it cost — are the whole decision.

What it hinges on. Two facts, neither public yet. One: does the proof survive formalization and independent review? Two: was it one shot or a search? Until both are answered, this is a data point, not a trend. What the council leans toward: the direction is real — AI math reasoning is climbing fast and the IMO-perfect-scores milestone confirms that — but this specific claim deserves the Skeptic's posture, not the headline's. If you build math-adjacent tools, rerun your evals now; whatever ceiling you measured six months ago is stale. Just don't ship anything that assumes the ceiling is gone.

Prediction: The Jacobian conjecture disproof will NOT be confirmed by a formalized, peer-reviewed proof (or a machine-checked Lean/Coq formalization accepted by the math community) within 90 days of the tweet — by roughly late October 2026.

Confidence: Medium — formal verification of pure-math results takes months, and the May Erdős result is still unresolved.

Why: Formalizing a novel negative result in algebraic geometry into a machine-checkable proof, or getting it through even preprint-level peer scrutiny, historically takes many months to years — the community is careful precisely because AI proofs have looked right and been wrong before. The May OpenAI Erdős disproof was announced with similar fanfare and is still working through scrutiny, which is the closest track-record we have and it points to slow. For this to be confirmed inside 90 days, the proof would have to be short, clean, and immediately formalizable — possible, but the base rate for 87-year-old open problems is against it. The opposite outcome — fast confirmation — would require the mathematics community to move faster than it ever has on a result this consequential, which is the less likely bet.

Revisit by 2026-10-24: We're right if there's no accepted formalization or peer-reviewed confirmation by then and the claim is still "under review" or quietly walked back. We're wrong if a machine-checked proof or a credible formal review confirms the disproof within the window.

Comments