Industry story
Anthropic's Claude Disproves 1939 Jacobian Conjecture During World Cup Final
evals interpretability reasoning
Claude apparently disproved the Jacobian conjecture, an 87-year-old open problem, mid-World Cup, in what an Anthropic researcher described as a chat with his "close friend Fable." The mathematics community is buzzing. Imperial College's Kevin Buzzard called it "a big day," and Demis Hassabis's May caution that AI breakthroughs are "far, far from what a Ramanujan would have done" is getting a fresh stress test. But the proof is a tweet, not a paper. No arXiv, no Lean formalization, no peer review. And the May Erdős disproof from OpenAI is still working through scrutiny. Two questions decide everything: does the proof survive formalization, and did it cost one clean session or ten thousand failed attempts?
Full analysis
An Anthropic researcher tweeted that Claude disproved the Jacobian conjecture, an 87-year-old math problem, apparently while he half-watched a soccer game. That's the story. The question for anyone building with these models: is math-grade reasoning now a shipping capability, or is this a viral tweet running ahead of peer review?
Reversibility: Type 2 for you. Nobody's signing a contract off a tweet. This is a "when do I rerun my evals" decision, not a "bet the roadmap" one. The mistake would be treating an unverified announcement as a settled fact.
What's actually being decided: Not "did Claude solve it." That's for the mathematics community to confirm. What you're deciding is whether frontier reasoning has crossed a threshold that should change how you benchmark, and how much of your own stack you trust to a model that can now allegedly do this.
Forcing function: None real. No deprecation, no eval window. Just narrative velocity, which is the weakest forcing function there is.
The Skeptic. The proof is a tweet. Read it again: "the Jacobian conjecture is false, thanks to my close friend Fable." That's not a paper. It's not on arXiv. It hasn't been formalized in Lean or Coq. One Imperial College professor, Kevin Buzzard, said "a big day," and Buzzard is genuinely excited about AI in math, so that's a friendly witness, not a jury. LLMs produce proofs that look airtight and collapse on line 40. The May Erdős disproof from OpenAI is still working through scrutiny. Announcement speed is outrunning verification speed, and each cycle makes the next tweet feel more inevitable. Plain version: someone claimed a world record but nobody's checked the tape yet.
The Safety Lens. Set aside whether it holds. If a conversational session can even plausibly attack an 87-year-old open problem, that's the capability cluster that worries alignment people: autonomous novel reasoning with no human in the loop for most of the chain. The "my close friend Fable" framing is charming and exactly the problem: it anthropomorphizes away the question of what the model did across the parts of the reasoning the researcher never read. For a builder, the lesson is narrower. If your model can invent a novel argument in math, it can invent a novel argument for why your guardrail doesn't apply. Interpretability has to run at the speed of capability, and right now it doesn't.
The Researcher. If it holds (big if), this matters more than the IMO scores everyone cited last year. The International Math Olympiad is hard, but the answers exist; the model is reproducing human-reachable reasoning under time pressure. The Jacobian conjecture is open. A genuine novel negative result on a problem that beat serious algebraic geometers for 87 years is a different tier than "assisted" or "verified." Demis Hassabis cautioned in May that disproving Erdős problems is "far, far from what a Ramanujan would have done," and that's the right prior. Inventing the right question is still the human's job. But the offhand, mid-match delivery tells you these researchers now run frontier models the way a 1990s mathematician ran Mathematica: ambient, habitual, background. That shift is real whether or not this specific proof survives.
The Compute Pragmatist. Here's the number nobody tweeted: the token budget. Was this one clean session, or ten thousand failed attempts and one lucky trajectory that got screenshotted? Those are wildly different capability stories and wildly different cost stories. If it's genuinely one long conversational session, the takeaway for your inference bill is encouraging. You don't need a giant multi-agent scaffolded research cluster; you need long-context high-quality reasoning at chat scale. If it's a needle in a haystack of retries, then "Claude disproved it" really means "a search process that happened to include Claude disproved it," and your per-result cost is enormous. Plain version: a slot machine that pays out once is not the same product as a machine that pays out on demand.
Where the council splits. The Researcher and the Skeptic aren't arguing about the model. They're arguing about evidence standards. The Researcher says the ambient, casual use is itself the signal; the Skeptic says casual delivery is exactly why you shouldn't trust it yet. The second split is between the Compute Pragmatist and everyone treating this as a clean capability jump: if the result took thousands of hidden attempts, the "conversational reasoning breakthrough" story dissolves into a "expensive brute-force search" story, and the product implications flip. Those two unknowns (does the proof hold, and what did it cost) are the whole decision.
What it hinges on. Two facts, neither public yet. One: does the proof survive formalization and independent review? Two: was it one shot or a search? Until both are answered, this is a data point, not a trend. What the council leans toward: the direction is real. AI math reasoning is climbing fast and the IMO-perfect-scores milestone confirms that. But this specific claim deserves the Skeptic's posture, not the headline's. If you build math-adjacent tools, rerun your evals now; whatever ceiling you measured six months ago is stale. Just don't ship anything that assumes the ceiling is gone.
Prediction: The Jacobian conjecture disproof will NOT be confirmed by a formalized, peer-reviewed proof (or a machine-checked Lean/Coq formalization accepted by the math community) within 90 days of the tweet, by roughly late October 2026.
Confidence: Medium. Formal verification of pure-math results takes months, and the May Erdős result is still unresolved.
Why: Formalizing a novel negative result in algebraic geometry into a machine-checkable proof, or getting it through even preprint-level peer scrutiny, historically takes many months to years. The community is careful precisely because AI proofs have looked right and been wrong before. For this to be confirmed inside 90 days, the proof would have to be short, clean, and immediately formalizable. Possible, but the base rate for 87-year-old open problems is against it. Fast confirmation would require the mathematics community to move faster than it ever has on a result this consequential, which is the less likely bet.
Revisit by 2026-10-24: We're right if there's no accepted formalization or peer-reviewed confirmation by then and the claim is still "under review" or quietly walked back. We're wrong if a machine-checked proof or a credible formal review confirms the disproof within the window.
Comments