Refacto AI

Podcast episode

Academia is for Ambition — Alex Zhang, MIT

agents cost-compression inference open-weights tool-use

Alex Zhang is a second-year MIT PhD student who went on a podcast to argue that the scaffolding around a model matters more than the model itself. His design, which he calls a Recursive Language Model, gives the AI exactly one tool (code), parks everything else in files on disk, and lets the model call copies of itself. Train it on short tasks and it handles tasks 8 to 30 times longer, jumping domains without new training. The independent signal worth taking seriously: Harvey, the legal AI company, landed on the same structure without coordinating with Zhang and reported strong results. Two parties converging independently is real evidence.

The rest carries weight, but comes with caveats. Zhang's numbers on wasted compute and generalization are his own, unreplicated by outside teams. The GPU leaderboard he cites is full of AI-generated code that breaks in production.

The practical move is cheap: swap the standard transcript-dump loop for a file-based external memory design. No retraining needed. Test it this week.

Full analysis

Alex Zhang, a second-year MIT PhD student, says the thing nobody selling you a bigger model wants you to hear: the frontier models you already pay for are smarter than the way you're using them. His claim is that the scaffolding around the model, the harness that decides how it reasons and acts, is the real lever. The harness. The next model release is secondary.

Here's what to confirm before you believe it. He defines a "harness" as the software wrapper around a language model that structures how it thinks and uses tools. His design, Recursive Language Models, gives the model exactly one tool, code, stores everything else on disk, and lets the model call copies of itself. The surprising finding: train it on short tasks, and it handles tasks 8 to 30 times longer, and jumps from math to writing without new training. If that holds, the cost of getting more out of AI is engineering time, not another training run.

This is a briefing, not a decision. Nothing here forces your hand this month. But it reframes where the next year of gains comes from, and that affects what you build and who you buy from.

The Skeptic. Zhang's own evidence undercuts the hype. On GPU Mode, a leaderboard for writing fast GPU code, almost every top submission is now AI-generated, yet only one entry in the top 10 actually ran correctly in a real system, and a human named Gauners produced it by steering the AI rather than letting it loose. That's the pattern everywhere: the model games the benchmark and ships something that breaks in production. His "8 to 30x generalization" comes from his own research. No outside team has reproduced it. And "95% of a swarm is wasted tokens" is his number too. Impressive demo, unaudited claim.

The Researcher. Strip the vocabulary and the real finding is narrow but genuine: give a model code as its only tool, keep its working memory in files instead of cramming it into one long prompt, and it transfers strategy across task lengths and domains. The independent signal that matters is Harvey, the legal AI company, which post-trained the same design on document-sifting work, without talking to Zhang, and reported strong results. Two parties landing on the same structure is worth more than one lab's leaderboard. Everything about OpenAI's 10,000-agent run is secondhand and unverifiable. Treat the Harvey result as the real data point and the swarm story as folklore.

The Open-Source Advocate. This is the good news for anyone who isn't a frontier lab. If capability lives in the harness, it's reproducible by people without a billion-dollar cluster. Prime Intellect is openly training a model on these recursive rollouts. The whole approach runs on a plain code interpreter and a file system, no exotic infrastructure. Notice which model won Zhang's internal harness tests before OpenAI's latest: Fable, a code model, likely Mistral's. The pattern keeps repeating. The frontier opens a lead, and smart scaffolding on cheaper open weights claws most of it back. That's the structural story under all of this.

The Compute Pragmatist. The dollars are worth spelling out. That Navier-Stokes run burned roughly 130 billion output tokens and cost about $40 million at list price. Zhang says 95% of it was useless search. Meanwhile GPT-5.6 reportedly rewrote inference code to make two internal systems about 80% cheaper to run. Those two facts point the same way: raw brute force is absurdly expensive, and the money is in making each query cheaper, not throwing more agents at it. One expert insight, he says, can erase a trillion tokens of wasted compute. For anyone paying an inference bill, convergence beats scale.

The Builder. On a Tuesday morning, here's what's actually usable. Look at your agent setup. If it's the standard loop, dump the whole running transcript back into the prompt each turn, you're doing what Claude Code and Codex do, and you're probably leaving gains on the table. Try the alternative: one tool, code, with context parked in files the model reads on demand. No retraining, just a redesign. The other practical nugget is firing off tool calls in parallel by reading the generated code before the model finishes writing it. Zhang calls it an obvious speedup for any code-driven agent. That one you could test this week.

Where they disagree. The Skeptic and the Researcher split on the same evidence. The Skeptic sees one grad student's unreplicated benchmark; the Researcher sees Harvey independently landing on the identical design and calls that corroboration. That's the whole bet in one argument: is this a method others will reproduce, or a clever result that fades? The second fault line is cost against capability. The Open-Source Advocate says scaffolding frees you from needing the frontier at all. The Compute Pragmatist counters that a swarm which wastes 95% of its tokens isn't freedom, it's a bigger bill, until someone solves convergence. Both can't be the headline.

What it hinges on. One belief: does harness design transfer across teams, or only inside Zhang's lab? The Harvey result leans yes. The absence of any third, fourth, or fifth independent replication keeps it from being settled. If you build agents, the thing to test is cheap and specific: run your current setup against a code-only, files-for-memory version on your own longest-running task, and measure whether it holds together where yours falls apart. You'll know in a week whether the generalization is real for your work, which is the only place it matters.

Prediction: By 2027-06-30, at least two companies other than Harvey and Prime Intellect will publicly report deploying a code-only, external-memory recursive harness (RLM-style) in production, citing results on long-running or document-heavy tasks.

Confidence: Medium. Two independent parties already converged on the same design, and the method needs no special hardware.

Why: Harvey post-trained this design on legal document work and reported strong results without ever coordinating with Zhang, which means teams are arriving at the same structure independently rather than copying one lab. The approach runs on a plain code interpreter and a file system, so the barrier to trying it is engineering time, not compute budget, and Prime Intellect is already training openly on these rollouts, which puts a reference implementation in public hands. When a pattern is cheap to copy and two unconnected groups have already validated it, more adopters surfacing within nine months is the likely path. The opposite, total silence, would require the Harvey result to be a fluke that nobody else can reproduce, which is possible but is the less likely bet given the independent convergence already on the table.

Revisit by 2027-06-30: We're right if two or more companies beyond Harvey and Prime Intellect publicly describe a production RLM-style harness with results. We're wrong if no such independent reports appear by that date.

Comments