Refacto AI

Podcast episode

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham - #778

agents evals gpu-supply inference open-weights

Greg Burnham runs capabilities research at Epoch AI, which tracks how fast AI models actually improve. He joined Sam Charrington on TWIML to make a claim that should reframe how you read every launch announcement: stitch the benchmarks across model generations and you get a straight line, with one small upward kink in late 2024 when reasoning models arrived. The "step-function breakthrough" is a marketing description of an on-trend result.

Two specific findings carry the weight. Burnham's team had models play an obscure board game repeatedly within a single session; GPT-5.6 scores the same on play one as play twenty. Models improve between generations, never during one. Any agent pitch promising a system that "learns your workflow" is overselling. The Navier-Stokes solve (one of math's seven famous unsolved problems) is real, but it required thousands of model copies running in parallel. Brute force, not a new idea.

Build accordingly: pin your model version, write your own memory layer, and plan for a step up when the next generation ships, then flat until the one after.

Full analysis

Greg Burnham runs capabilities research at Epoch AI, an outfit that measures how fast models improve. His headline claim on the TWIML podcast: every splashy release you've read about, GPT-6, Google's Astra, whatever came next, lands exactly where the trend line already predicted. The "step-function breakthrough" is marketing. What moves underneath is a straight line, and it hasn't slowed.

That matters because you buy AI on the strength of vendor claims. If the trend is smooth and predictable, so is your planning. Two things could break that: a machine that solves genuinely new problems, and a machine that gets better mid-task without a retrain. Burnham says neither has happened yet. Both are being watched.

How hard is this to undo? Nothing here forces a decision. This is a read on the pace of the field, useful for how you budget and how much you trust the next launch deck. No deadline, no contract, no shutdown date. Low urgency, high value as a mental correction.

The Skeptic

The useful debunk here is the "revolutionary release" story. Burnham has the receipts: stitch the benchmarks together across generations and the line is straight, with one small upward bend in late 2024 when reasoning models arrived. So when a vendor tells you the new model is a leap, translate that yourself: it's on trend, and the trend is fast. Fine. But watch what Burnham quietly concedes. The Navier-Stokes solve, one of the seven famous unsolved math problems, took a swarm of thousands of model copies grinding in parallel. That is not intelligence. That is brute force with a big electricity bill. Impressive, expensive, and not the same thing as a new idea.

The Researcher

Two gaps do the real work in this conversation. First, no model has produced math's version of a genuinely new move. Fields Medalist Timothy Gowers looked at the Erdős problem the AI cracked and judged that a human with a two-part hint would have gotten there too. The machine redirected the search. It did not invent the road. Second, the on-the-fly learning test: Epoch had models play an obscure board game, Earthborn Rangers, over and over. GPT-5.6 scores the same on play one and play twenty. GPT-6 scores higher, then also flatlines. Models improve between generations, never during a session. That is the ceiling on any agent you deploy expecting it to learn the job.

The Builder

This is the line that should change how you scope agent projects. If a model can't get better within a session, then every "autonomous agent that learns your workflow" pitch is overselling. What you actually get is a fixed skill level per model version, plus whatever you can cram into its working memory with notes and prompting. Burnham calls those gains low-hanging fruit. So build accordingly: pin the model version, write your own memory layer, and don't budget for a system that compounds on the job. It won't. When GPT-6 ships you get a step up, and then flat again until the next one.

The Compute Pragmatist

The Navier-Stokes result is compute bought as a trophy. Thousands of instances running in parallel to solve one problem means the frontier's most dazzling results are bought with compute, not cleverness. For anyone renting GPUs, that sets the price of "genuinely hard reasoning" very high. The other half matters more for the whole field. Burnham's tripwire eval asks whether a model can rediscover a recent AI research advance on its own, without being told what it is. No model has passed. The day one does, GPU supply stops being the thing that gates progress, because the machine starts improving the machine. Until then, chips and power remain the binding constraint, which is oddly reassuring for planning.

The Open-Source Advocate

One thing the closed labs can't hide inside a benchmark: the smoothness of the curve means capability is a commodity moving on a schedule everyone can see. Epoch measures it across vendors and finds no secret sauce, just position on a line. That is good news for open weights. If there's no magic leap, only steady grind, the gap between the best closed model and a good open one is a matter of months and money, not a moat. The exception is the swarm-based brute-force stuff. Running thousands of instances is a rich-lab move. Open models can match the reasoning; matching the compute budget behind a Millennium Prize attempt is another thing entirely.

The tensions

Three real disagreements sit under this.

The Skeptic and the Researcher both say "no new ideas yet," but they read the risk differently. The Skeptic hears brute force and relaxes. The Researcher notes the tripwire hasn't been tripped yet, and that the people inside the labs are, in Burnham's word, spooked. Same facts, opposite comfort levels.

The Compute Pragmatist and the Open-Source Advocate part ways on the moat. The Pragmatist sees compute as the wall that keeps open models out of the top tier. The Advocate sees a straight, public trend line and concludes the wall is low and getting lower. Both are right about different layers: reasoning quality commoditizes, swarm-scale compute does not.

And the Builder against everyone selling agents. The whole "learns on the job" category runs into a flat line in Epoch's data. If Burnham is right, a lot of agent roadmaps are priced on a capability that doesn't exist yet.

What this hinges on

Two facts, and Burnham names both. Can a model generate a genuinely new idea, not just search harder? And can it improve within a session? Both are "no" today, measured, not vibed. Your planning should assume they stay "no" for now, because Epoch is watching the exact tripwires that would flip them and reports nothing has moved. If either flips, the smooth trend line bends and everyone's timelines compress. Until then, buy AI as a fast-improving but predictable commodity: pin your versions, build your own memory, and discount any "learns and adapts on its own" claim to zero.

Prediction: Through the next full model generation from OpenAI, Google DeepMind, or Anthropic, Epoch AI's in-context learning test (repeated plays of a game like Earthborn Rangers) will still show flat within-session scores, with improvement only between generations, as of Epoch's next capabilities report.

Confidence: Medium. The gap is fundamental to how these models are built, not a tuning issue.

Why: Burnham's data shows every model to date, GPT-5.6 and GPT-6 included, scores the same on the first and twentieth play of a novel game, because the weights are frozen after training and the working memory can't substitute for real learning. This is an architecture limit, not a scale limit, so a bigger or better-trained model of the same kind lands in the same place, exactly as Epoch's straight trend line predicts. The opposite result, a model that measurably improves across a single session, would require a genuinely new training method that no lab has shipped or announced, and Burnham, who runs the eval, reports no sign of it. The safe bet is that the next generation is smarter out of the box and still flat within a session.

Revisit by 2027-04-04: We're right if Epoch's next capabilities work still reports no within-session improvement on its on-the-fly learning evals for the newest frontier models. We're wrong if any of OpenAI, Google DeepMind, or Anthropic ships a model that Epoch measures improving across repeated attempts inside one session.

The tripwire is the one to hold in reserve. Burnham built the "can it rediscover a recent AI advance on its own" test precisely because the day it passes, the compute-gated, predictable world he describes stops being predictable. He says nothing has passed. Believe that until Epoch says otherwise, and treat anyone claiming self-improving AI today as selling the thing the people who measure it can't find.

Comments