Refacto AI

Podcast episode

Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776

cost-compression evals inference model-pricing open-weights

Stanford linguist Chris Potts joined Sam Charrington on the TWIML AI Podcast to make a case that should get anyone paying for AI coding tools out of their chair: the value of a token is falling, even as the models get smarter. Potts calls it "tokenflation," and he built a price index for AI the way the government builds one for groceries, running it across roughly 6,000 real coding sessions to measure it.

The finding that sticks: a fixed model produced wildly different economics depending on product settings, reasoning depth, context length, system prompt. The leaderboard score tells you almost nothing about your actual bill. Potts also found that expert users who push back and iterate extract far more value than novices who accept the first answer, which means broad rollouts to non-expert staff will underperform whatever your internal champions promised.

One benchmark, one model, ten weeks. This needs replication before "tokenflation" is a trend. But instrument tokens-per-outcome now, not tokens-per-call.

Full analysis

Here's the claim that should stop you mid-scroll: the tokens you buy from an AI coding tool are getting worse value, month over month, even after you account for the models getting smarter. Stanford's Chris Potts calls it "tokenflation." He built a price index for AI the way the government builds one for groceries, ran it on about 6,000 real coding sessions, and found the purchasing power of a token falling from February to mid-April 2025 on Claude Code running Opus 4.6.

If you're buying AI by the token, or paying per-seat for tools that burn tokens on your behalf, that's your bill going up while your output stands still.

How hard is this to undo? For you as a buyer, easy. You can instrument your spend, switch tools, or cap reasoning depth this quarter. For the labs, hard: the plateau Potts describes is baked into the architecture. What's actually being decided is whether operators keep treating benchmark scores as the buying signal, or start watching cost-per-outcome. Nothing sets a hard deadline, but the direction of travel is one-way: prices are moving toward true cost as the big labs prep for IPOs.

The Skeptic. One benchmark, one model, one ten-week window. Potts measured "code survival," lines that stay in the repo more than four days, on Anthropic's Opus 4.6. And he admits the experiment got confounded when Anthropic changed the default reasoning setting mid-study, from high to medium and back. So part of his "tokenflation" is Anthropic quietly turning up how much the model thinks, then billing you for the thinking. That's a product knob, not a law of nature. Turn adaptive reasoning off and the effect may shrink. Before anyone declares an era of declining token value, I want this replicated on GPT, on Gemini, on a task that isn't coding. Right now it's a striking chart, not a trend.

The Researcher. The deeper point survives the small sample. Potts is saying the flattening of inference-time scaling, spending more "thinking" tokens for smaller and smaller gains, was already predicted in the literature, and the token bloat we see now is that prediction showing up in your invoice. His architecture read is the part worth chewing on: today's models have fixed depth and no real recursion, so the only way to buy more computation is to generate more tokens. That's the mechanism behind diminishing returns. He also notes the 2017 transformer got re-engineered by people who understood the math, not by dumb scaling, which cuts against the "just add compute" story. If he's right, the cost curve doesn't bend until the architecture changes.

The Builder. The line that's immediately useful: a fixed model produced wildly different economics depending on product knobs, reasoning depth, context length, system prompt. So the leaderboard tells you almost nothing about your bill. Instrument tokens-to-outcome now. Not tokens per call, tokens per thing that shipped and stuck. Potts's other finding maps straight onto how you roll AI out to a team: expert users who push back, iterate, and complain get more out of the model; novices who delegate and accept the first answer plateau. That means broad rollout to non-expert staff will underperform whatever your internal champions promised. The lever is UX and training that scaffolds pushback, not another model swap.

The Compute Pragmatist. Potts pegs the true cost of a token at 2x to 20x what you pay today, with the spread that wide because nobody outside the labs knows how R&D gets amortized. He points to the viral Copilot bill that jumped from about $500 to a projected $11,000 under new terms, and the ride-share analogy: same trip, $20 becomes $500. The subsidy is ending, and IPO prep is what's ending it. For anyone whose unit economics assume today's token price holds, that's the number that breaks the model. And the architecture escape hatch, recursive designs, byte-level models that drop the tokenizer, is real research but not shipping in your procurement window.

The Open-Source Advocate. The quiet argument here is that you need the "weird players" to stay alive. Potts leans on open-weights models as the thing that keeps architecture diverse, and flags byte-level work (his student Julie Colini's) as underexplored. If the frontier labs all converge on the same expensive stacked-transformer recipe and then reprice it toward true cost, the pressure valve is an open model that's within reach on a task you care about, running on hardware you control. Mixture-of-experts, where only a slice of the model fires per token, already broke one set of scaling assumptions. The next disruption to your GPU bill is more likely to come from an architecture change than from a price cut by a lab heading into an IPO.

Where they part ways

Is tokenflation a law or a knob? The Skeptic says Anthropic turned up reasoning and billed you, so fix the setting and the problem shrinks. The Researcher says the setting is downstream of an architecture that has no cheaper way to compute, so the knob is a symptom. This is the whole disagreement. If it's a knob, you manage it in config. If it's the architecture, you manage it in your budget for years.

Do benchmarks still guide buying? The Builder says no, the system-level knobs drive your bill and the leaderboard hides them. The Researcher still wants the benchmark, just an economic one, cost-per-surviving-output, run on your own traffic. Both agree the MMLU-style score on a slide is the wrong number to sign a contract on.

Does relief come from the labs or from outside them? The Compute Pragmatist watches the frontier labs raise prices into their IPOs. The Open-Source Advocate says that's exactly why the open ecosystem matters, because the cheaper architecture won't come from the incumbent selling you the expensive one.

What it hinges on

Three beliefs. One: is the declining token value real beyond Opus 4.6 and coding, or an artifact of one lab's reasoning default? Two: does the inference-cost plateau force the big labs to pass through real costs before a cheaper architecture arrives? Three: does expert-style prompting actually drive ROI enough to justify training spend? The council leans toward the plateau being real and the repricing being underway, and toward the fluency gap being a genuine and underinvested lever.

What to do before committing anything: instrument cost-per-outcome on your own workload this month, not tokens per call. Test whether capping reasoning depth changes your bill without hurting your results, that's a one-afternoon experiment. And if a vendor contract assumes today's token price, ask for price-protection language, because the labs are telling you plainly that today's price is a subsidy.

The call

Prediction: By June 30, 2027, at least one of the three major coding-AI providers (Anthropic's Claude Code, OpenAI's Codex/GPT coding tools, or GitHub Copilot) will publicly raise per-token or per-seat pricing, or tighten usage limits at existing prices, on its top coding tier, citing inference cost.

Confidence: Medium. The subsidy-ending mechanism is clear; timing and who moves first is not.

Why: Potts estimates true cost-per-token runs 2x to 20x current pricing, and ties the repricing directly to IPO preparation at the large labs, who can't run permanently subsidized inference into a public offering. The Copilot bill jumping from ~$500 to a projected $11,000 under new terms is that pass-through already starting. The mechanism is simple: a lab preparing for public markets has to show a path to positive unit economics, and coding agents are the heaviest token burners, so they're where the subsidy hurts most and gets cut first. The opposite outcome, prices holding or falling across all three, would require the plateau in inference-time scaling to break, or a cheaper architecture to ship at scale, and Potts's argument is precisely that neither is close.

Revisit by 2027-06-30: We're right if any of Anthropic, OpenAI, or GitHub raises headline price or cuts included usage on a top coding tier and points to cost. We're wrong if all three hold or cut effective price on those tiers with no new usage caps through that date.

Comments