Refacto AI

Podcast episode

Why smarter AI models could drive up compute prices 10x

cost-compression gpu-supply inference model-pricing open-weights

Dwarkesh Patel ran the arithmetic on frontier AI compute and came up with a number that should make anyone running AI in production uncomfortable: if models reach human-level software engineering capability, an H100-equivalent GPU (today's standard training chip, renting for roughly $17K a year at spot) should rationally fetch $250K a year. His mechanism is a straightforward supply-demand squeeze. Anthropic's revenue has grown 10x annually for three years running, while global GPU supply grows only about 3x a year. He also points to a physical ceiling on wafer supply he dates to end of 2026, after which even that 3x growth rate is in question.

The most interesting counterargument is the one Patel raises himself: this pattern-matches to every commodity scarcity bet that lost. A lot of that demand is subsidized or free-tier usage that evaporates the moment prices move.

For an ad-tech engineering lead, the practical move is simpler than the macro argument. Separate your latency-critical calls from your batch generative workloads, and make sure the batch tier can fall back to an open-weight model on owned capacity. The AI features whose business case only works at today's token price are the fragile ones.

Full analysis

Dwarkesh Patel ran the arithmetic on frontier compute and landed on an uncomfortable number: if AI hits human-level software engineering, an H100-equivalent GPU should rationally rent for ~$250K a year, roughly 15x today's spot price. His logic is a supply-demand squeeze. Anthropic's revenue has 10x-ed three years running while global compute only grows ~3x a year, and the two ways out (fatter lab margins or higher compute prices) are running out of room on the margin side.

What's actually being decided: whether teams shipping AI into production should treat cheap inference as a temporary subsidy and plan for a world where tokens get materially more expensive. This is a Type 2 call for most reads (you can renegotiate contracts, swap providers, re-architect), but the long-term compute contracts Patel points at are Type 1. Forcing function: the wafer-absorption ceiling he dates to end of 2026. Below, the council pressure-tests the thesis and what an ad-tech engineering lead should do with it.

The Skeptic. Patel's own caveat is the strongest line in the essay: this pattern-matches to every scarcity bet that lost, and he name-checks Simon-Ehrlich to prove he knows it. The tell is the revenue extrapolation. "$1T by end of 2026 if the trend holds" is doing all the work, and three data points of 10x growth is a slope you drew through a startup's early ramp, not a confirmed trend. Revenue 10x-ing does not mean paid demand 10x-es; a lot of that is discounted, subsidized, or free-tier usage that evaporates the moment prices move. For a PM: the scary number assumes AI keeps getting more useful and everyone keeps paying full freight. Break either and the squeeze softens.

The Researcher. Start with the inference split, not the GPU valuation. OpenAI ran ~25% of compute on inference in 2024, ~50% now, per Epoch AI. That shift is real and measurable, and Patel's read on why labs resist tilting further is the interesting bit: more inference share signals training progress stalled, which undercuts the AGI story they raise money on. So the compute allocation is partly a financing decision, not a physics one. The $250K/GPU figure, by contrast, is a thought experiment resting on "true human-level software engineer on one H100," which no benchmark today supports. SWE-bench gains are real but nowhere near "replace a $250K engineer end to end."

The Open-Source Advocate. Every line of this thesis assumes you're renting frontier tokens from a lab that can outbid you. That's the escape hatch. If compute gets 10x more precious, the pressure to run Qwen, Llama, or a distilled open model on your own reserved capacity goes way up, and for most ad-tech workloads (creative gen, classification, bid-adjacent scoring) an open model at 80% of frontier quality already does the job. Patel's own "Alkan Allen Effect," that the most token-efficient model wins when compute is dear, cuts toward small efficient open weights as much as toward frontier labs. The squeeze he describes is a reason to own your inference, not rent it.

The Compute Pragmatist. The price signals are the part I'd act on. Google paying $900M/month for 110,000 GPUs at 2x spot, and spot itself already 40%+ above the February 2025 trough, tells you reservation premiums are widening right now, independent of whether AGI ever shows up. That's not a forecast, that's a print. The wafer ceiling is the constraint with teeth: 3x annual supply decomposes into ~1.4x Moore's Law, ~1.2x new fabs (gated by ASML EUV output through 2030), and ~1.8x from AI eating smartphone and PC wafer share. That last 1.8x dies when AI hits ~86% of leading-edge wafers, projected end of 2026. After that the 3x itself is in question.

The Builder. What do I do Tuesday? Not much that's dramatic, and that's the point. The consumer/commodity squeeze is the real operational risk: if labs would rather sell tokens to agentic research than to your bulk creative-variation pipeline, the batch jobs where you burn millions of cheap tokens are exactly what gets repriced. So I'd instrument cost-per-outcome per workload today, separate the latency-critical RTB-adjacent calls from the batch generative stuff, and make sure the batch tier can fall back to an open model on owned capacity. For a PM: the AI features whose business case only works at today's token price are the fragile ones. Find them now.

Where the council splits. Three real disagreements. First, the Skeptic and the Compute Pragmatist part ways on why prices rise: the Pragmatist sees a physical supply cap that holds regardless of demand hype, the Skeptic sees a demand curve propped up by subsidized usage that could snap. You can believe prices rise for the first reason and not the second. Second, the Researcher and the whole $250K thesis: the valuation only lands if models actually reach human-level engineering, and the evals don't say that yet. Third, the Open-Source Advocate versus the frontier-lock assumption: Patel's math only bites teams who must rent frontier tokens, and most ad-tech inference doesn't need them.

What it hinges on. Two beliefs. One, does compute supply growth actually stall near the wafer ceiling on Patel's timeline. That one has hard, observable inputs (ASML EUV shipments, TSMC leading-edge allocation) and looks more solid than the demand side. Two, does frontier capability keep justifying frontier prices, or does the open-weight tier close enough of the gap that the squeeze routes around you. For an ad-tech shop, the second is the one you control. The council leans toward "prices go up, but you have an exit": lock reserved capacity for latency-critical work, move batch generative to owned open-model inference, and stop building features whose margin only survives at 2025 token prices.

What to verify before signing anything long: benchmark your top three workloads at output-quality-per-dollar on both a frontier API and a hosted open model, and negotiate any multi-year GPU contract with a repricing or exit clause, because you're contracting into a market Patel himself admits could break either way.

Prediction: Frontier inference list prices (per million tokens on OpenAI's and Anthropic's flagship models) will NOT rise 10x by end of 2026; the headline price of the leading general model will be flat or lower than August 2026 levels when the wafer-ceiling milestone Patel dates to end of 2026 arrives.

Confidence: Medium. Competition and efficiency gains have cut per-token prices every year so far.

Why: The thesis needs demand to keep 10x-ing while supply caps out, but the observable trend on published token prices runs the other way: every major model generation since GPT-4 has shipped at a lower per-token price than the one before, driven by distillation, MoE routing, and quantization, not by more wafers. Patel's own "efficient model wins" effect pushes labs to compete on tokens-per-dollar, which shows up to buyers as flat or falling list prices even if raw GPU rental costs rise. The 10x compute-price squeeze can be true at the bare-metal GPU layer while frontier API prices stay flat, because labs eat the gap in margin and efficiency rather than pass a 10x to customers who would defect to open weights. Prices spiking 10x would require the leading lab to have no efficiency lever left and no competitor, which nothing in 2026 supports.

Revisit by 2026-12-31: We're right if OpenAI's and Anthropic's flagship per-million-token prices are at or below August 2026 levels. We're wrong if either lab's flagship list price is 2x or more above August 2026.

Comments