Industry story
OpenAI AI solves 100+ open math problems, forms advisory group
OpenAI claims an internal model has solved more than 100 open math problems, including Navier-Stokes, and stood up a nine-person advisory group at Princeton's Institute for Advanced Study to manage the academic blowback. The advisory group can help schedule announcements but cannot touch the pace of research, and IAS was careful to note it has no decision-making power over OpenAI. That structure tells you exactly what this is: borrowed prestige with no accountability attached. Until these results exist as machine-verified proofs in Lean or Coq, checked by mathematicians with no OpenAI affiliation, "solved" means output nobody has falsified yet.
Full analysis
OpenAI says an internal model has cracked more than 100 open math problems, on top of an earlier claim that it solved Navier-Stokes, one of the seven Millennium Prize problems. To calm a furious academic community, it's standing up an advisory group of nine mathematicians hosted at Princeton's Institute for Advanced Study. The catch, spelled out in the announcement: the group can help time the release of results, but cannot touch the pace of OpenAI's research, and IAS has no decision-making power over anything.
Here's what's actually being decided, and it isn't math. This is a governance move that's easy for OpenAI to undo, because the advisory group has no power to begin with. There's no deadline forcing anyone's hand. So the real question for people who buy and build with AI is narrow: does this tell you anything about what OpenAI's models can now do that you'll be able to use, or is it a claim you can't check?
The Skeptic. A "resolved" problem with no published proof, no reproducible steps, and no community verification is not a solved problem. It's output nobody has falsified yet. The Navier-Stokes claim alone should have drawn months of scrutiny before anyone announced 100 more. Instead the score jumped from one famous problem to "most areas of mathematics" in a single press cycle. That's the pattern of a marketing calendar, not a proof pipeline. My bet: a large share of these turn out partial, conditional on unproven assumptions, or quietly dropped once real mathematicians get the full write-ups. The advisory group is the PR wrapper for exactly that walk-back.
The Safety Lens. Twenty-five Fields Medalists signed an open letter, and OpenAI's response was to build a body that can schedule announcements but cannot slow the work. Read that plainly: the domain experts flagged a problem, and the fix removes their leverage while borrowing their prestige. IAS itself said it has no say over OpenAI. If a lab treats the people best placed to check its most consequential scientific claims as a communications issue to manage, that's the governance gap alignment researchers keep pointing at. The outputs look like math, so the risk feels contained. Foundational physics claims released without the machinery to know they're right is not contained.
The Researcher. If 100 open problems genuinely fell, this is the biggest scientific event in a generation. That's exactly why the absence of proof matters. Math has a verification process for a reason. A claim isn't a result until it survives adversarial reading by people trying to break it. Quality proof would look like this: results checked in Lean or Coq, the formal proof systems where a computer verifies every step, with the proofs published. Nothing in the announcement says that happened. Fields Medalist fury isn't technophobia. It's the people who own the standard saying the standard was skipped.
The Builder. Set the drama aside and ask what ships. If the model really generalizes across mathematical domains, the useful downstream work is formal verification and theorem proving, and better code-correctness tooling on top of that. But that only matters to me if the proofs are machine-checkable. If OpenAI can hand a mathematician a Lean file that verifies, it's engineering-usable. If it hands over prose, it's a claim. And there's no API. The distance between "our internal model solved it" and "I can call something that reliably proves things about my codebase" is enormous, and OpenAI said nothing to close it.
The Compute Pragmatist. Whatever ran here sits well above anything OpenAI has put in public hands. Either the internal frontier is much further ahead than the released models, or this used a compute budget no inference operator is pricing. Both readings are expensive. Every rival lab will read "math reasoning pays" and justify the next giant training run, which is more demand for NVIDIA's top chips. But none of that reaches your bill or your tools until there's a product, and there isn't one.
Where they split. The Compute Pragmatist and the Skeptic look at the same announcement and see opposite things. One sees an internal capability leap real enough to reset the training arms race. The other sees output that hasn't been proven and probably won't fully hold. The Builder cuts through both: it doesn't matter which is right until the proofs are machine-checkable and there's something to call. And the Safety Lens names the part everyone else steps around: the one group that could settle whether these results are real was handed a role that explicitly can't.
What this hinges on. One fact decides everything: are the proofs formally verified and published, or not? If they're checked in Lean or Coq and released, the Skeptic loses and this is historic. If they stay as prose that OpenAI "assessed" internally, the burden of proof hasn't moved, and the advisory group is there to manage the gap. Before you treat any of this as real capability you can lean on, wait for a proof a machine confirmed and a mathematician outside OpenAI vouched for. Neither exists today.
Prediction: Of the more than 100 open math problems OpenAI claims its internal model resolved, fewer than half will have full machine-verified proofs (in Lean, Coq, or an equivalent system) published and independently confirmed by a mathematician unaffiliated with OpenAI by the time the IAS advisory group issues its first public assessment, or by 2027-06-30, whichever comes first.
Confidence: Medium. The incentive to announce beats the incentive to verify.
Why: OpenAI moved from one claimed Millennium-problem solution to "more than 100 across most areas of mathematics" in a single announcement, with no published proofs and no mention of formal verification, which is the step that turns a claim into a result. It then built an advisory group that can time releases but explicitly cannot slow the research, and hosted it at IAS, which stated it has no decision-making power. That structure protects the announcement rather than the epistemics: it lets OpenAI keep claiming results while the slow work of verification lags far behind the press. The opposite outcome, most of these fully checked and confirmed within nine months, would require formal proofs of 100+ open problems, and formalizing even a single hard result in Lean routinely takes specialists months. The math can't be verified as fast as it was announced.
Revisit by 2027-06-30: We're right if fewer than half the claimed resolutions have machine-verified, published proofs confirmed by an unaffiliated mathematician by then. We're wrong if half or more do.
Comments