Refacto AI

Industry story

OpenAI AI Swarm Solves Millennium Prize Math Problem in 88 Hours

agents evals orchestration reliability

OpenAI deployed a 'swarm' — thousands of AI agents running in parallel — to tackle the Navier-Stokes existence and smoothness problem, one of the Clay Institute's seven Millennium Prize Problems (each worth $1 million). The swarm, coordinated loosely by OpenAI with agents passing the best ideas between groups via Codex, exchanged roughly 2.7 million messages and reached a result in 88 hours. The Clay Institute has not yet formally accepted the proof, but according to the author is treating the result as settled.

The author, Wharton professor Ethan Mollick, highlights that the coordination overhead was minimal: OpenAI set the goal and made one major directional shift, while the agents self-organized internally. This contradicts his earlier prediction that managing large agent swarms would require elaborate human-designed structures analogous to building a company. He attributes the surprise to 'The Bitter Lesson' — the recurring pattern in AI where brute-force scaling of machine learning outperforms hand-crafted human rules and systems.

Analysis

Showing the shorter version.

OpenAI says a swarm of thousands of AI agents solved the Navier-Stokes problem, one of seven unsolved Millennium Prize Problems worth $1 million each, in 88 hours. Wharton professor Ethan Mollick reported it and says the Clay Mathematics Institute, which administers the prize, treats the result as settled. The Institute itself has said nothing publicly.

Separate the two claims, because they have different odds of holding up.

The math is not settled. Millennium proofs take months to years of formal peer review, and the Clay Institute's own rules require a proof to appear in a peer-reviewed journal and then survive a two-year waiting period before a prize committee even convenes. Mollick's read is secondhand. Navier-Stokes has attracted wrong proofs before. The prediction here: the Clay Institute will not formally accept this proof by April 2027, six months out. High confidence.

The coordination result is the genuinely interesting thing, and it stands whether or not the proof holds. The swarm passed 2.7 million messages among itself, self-organized, and required one directional nudge from OpenAI across the full 88 hours. If your team has been building heavy orchestration for multi-agent work, assigning roles, routing, review gates, this says that scaffolding may be cheaper than you assumed. Worth a test on your own problems.

Two caveats. First, one run on one problem is an anecdote. Before rebuilding any pipeline around "set the goal, trust the swarm," run it on a problem where you already know the answer and can measure how much structure you actually stripped out before quality fell. Second, 88 hours of thousands of parallel agents is not cheap. OpenAI ran it on their own hardware, so the cost is invisible in this story.

The safety point is worth keeping in mind. Navier-Stokes is safe for an unmonitored swarm because the entire math community can check the output line by line. Extend the same setup to biotech or security research, where the answer cannot be audited by outside experts after the fact, and the risk profile is different. Don't generalize the comfort from this particular problem.

The council is skeptical on the math and genuinely interested in the coordination claim. When OpenAI or anyone else publishes a second hard-problem run with the same low overhead, that result is what tells you whether to touch your pipelines.

Also covered this issue

Comments