Refacto AI

Industry story

OpenAI AI Claims 700+ Math Solutions, Potentially Historic

cost-compression evals model-pricing reasoning

Commentator Alberto Romero writes that OpenAI has published a document containing over 700 mathematical solutions at varying stages of verification, many of which — if proven correct — would be of significant historical importance, with some potentially surpassing the importance of solving the Navier-Stokes equations (a centuries-old unsolved problem in fluid dynamics). Romero quotes serious observers calling this the 'biggest thing ever to happen in mathematical history' and notes that several Fields Medal-caliber results may be contained within it.

Romero frames this as a civilizational inflection point, drawing a parallel to AI's conquest of chess and Go: just as Magnus Carlsen can no longer beat or even fully understand Stockfish (a chess engine), humans may now be permanently outpaced in mathematics. He argues that OpenAI staffers encouraging people to 'learn the magic' are engaged in wishful thinking — that 'learning from superhuman intelligence is not the same thing as remaining intellectually commensurate with it' — and that the human relationship to intellectual disciplines like math is being fundamentally severed, even if the disciplines themselves survive.

Analysis

Showing the shorter version.

OpenAI released a document claiming 700-plus mathematical "solutions," and Alberto Romero called it a civilizational inflection point, comparing it to a nuclear bomb for mathematicians. TechCrunch, the same week, reported the solutions aren't meeting the field's standards yet. Both are in the public record. Figuring out which signal to trust is the whole exercise.

Start with what's missing. OpenAI hasn't said what actually produced these results. That single fact decides everything downstream. Machine-checkable formal proofs in Lean or Coq could be verified in weeks and would represent a genuine capability leap. Natural-language arguments needing human expert sign-off are a release strategy, and could drag on quietly for a year before deflating. Until OpenAI publishes the method, you're reading a marketing document.

"At varying stages of verification" covers a lot of ground. It could mean 695 are garbage and 5 are incremental. The prior on one system generating hundreds of correct novel proofs simultaneously is low, and Andrew Wiles's proof of Fermat's Last Theorem, far more scrutinized than anything in this drop, had a serious gap surface months later. That's the base rate for deep mathematical results.

The asymmetry matters, though. If even three to five hold up at the claimed level, this is genuinely historic, and the 695 wrong ones stop mattering. That's why the story won't die quietly.

One underpriced concern: if this is fully real, it means humans have lost the ability to check the work at the speed it's produced. Magnus Carlsen can't beat or fully follow a chess engine anymore. The same dynamic applied to mathematical reasoning means the "human in the loop" safety story stops working, because you can't supervise what you can't verify in time.

The release format is itself evidence. Genuine Fields-level work goes to journals and proof-checkers because the author wants the stamp. A document drop wants the headline first and the verification later.

The call: By April 2027, the number of OpenAI's claimed solutions independently confirmed by established mathematicians as novel, correct, and of the claimed historic importance will be in the single digits. Hundreds confirmed is not what the evidence supports, and the field has never once verified work at that speed.

Also covered this issue

Comments