Podcast episode
Why smarter AI models could drive up compute prices 10x
cost-compression gpu-supply inference model-pricing open-weights
TL;DR
Dwarkesh Patel works through the arithmetic of why frontier AI compute prices may rise 10x or more: Anthropic's revenue is 10x-ing annually while global compute supply only 3x-es, and the only escape valves are higher lab margins (already at 80%+, but probably can't sustain >90%) or rising compute prices. If AI reaches human-level software engineering capability, an H100-equivalent GPU should rationally rent for ~$250K/year — 15x today's spot price. This is a solo essay read, not an interview, but the macro argument is directly relevant to anyone sizing AI infrastructure bets.
What was covered
- Anthropic revenue trajectory: Anthropic has 10x-ed revenue three consecutive years. Ended 2024 at ~$9B; Patel projects $100B–$150B in 2025, and ~$1T by end of 2026 if the trend holds. Framed explicitly as "wild" and contingent on capability milestones.
- Compute supply constraint — the 3x ceiling: Global AI compute capacity grows roughly 3x per year, decomposed as: ~1.4x from Moore's Law, ~1.2x from new fab construction (bottlenecked by ASML EUV machine production through at least 2030), and ~1.8x from AI absorbing wafer share previously allocated to smartphones/PCs. That last factor hits a ceiling when AI reaches ~86% of leading-edge wafer capacity (projected by end of next year, up from 60% now).
- Inference vs. training compute split: Per Epoch AI data, OpenAI allocated only ~25% of compute to inference in 2024; that figure is now estimated near 50%. Labs resist tilting further toward inference because it signals that training progress has stalled — undermining the AGI narrative needed to raise capital.
- Inference margins: Anthropic's inference margins reportedly rose from ~40% (mid-2024) to 80%+ now. Patel argues margins above 90% are theoretically possible only if the leading model is so far ahead that competition is non-existent — which he finds implausible.
- Google/SpaceX compute deal as a price signal: Google is paying $900M/month for 110,000 GPUs (a blend of GB200s and GB300s) — reportedly 2x the current spot price. Spot prices themselves are already 40%+ above the February 2025 trough.
- Human-level AI and the implied GPU valuation: If an H100-equivalent could run a true human-level software engineer (who earns ~$250K/year in the market), that GPU should theoretically rent for $250K/year — 15x current spot. Patel invokes the "lump of labor fallacy" to argue that a massive AI labor supply shock may not necessarily collapse marginal compute value, per standard economic theory.
- Competitive moat for efficient models: Patel introduces what he calls the "Alkan Allen Effect" — when compute is expensive, the lab with the most token-efficient model can charge a large premium because users avoid burning costly compute on weaker models. Efficiency in model training becomes a structural competitive advantage.
- Current cheap AI apps may get priced out: As compute becomes more valuable, frontier labs will outbid consumer/commodity use cases for tokens — prioritizing high-value agentic workloads (e.g., AI-driven research) over low-value generative content.
Notable claims & predictions
- Patel: "I think [Anthropic] will probably end this year with somewhere between $100 billion to $150 billion dollars in revenue" — a 10x+ from the $9B reported for 2024.
- Patel: "If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers that H100 should rent for over $250K a year — that's over 15x the current spot price."
- Patel: "Google is paying $900 million a month for 110,000 GPUs… the price Google is paying is 2x the spot price per hour for those GPUs, and that spot price itself is more than 40% higher than it would have been in February."
- Patel: "At the leading edge… AI will have gone from 60% to 86% [of wafer allocation]. At some point you have just absorbed all leading-edge wafer capacity for AI and you can't keep increasing this number."
- Patel: "If you can train the best, most efficient model then you'll be able to charge much larger margins… if you have a model that can get the same result by using less compute, you've in some sense created more compute."
- Patel (caveat): "I'm a bit worried that this analysis honestly pattern-matches a lot onto ways people in the past have been wrong about scarcity" — citing the Simon-Ehrlich commodity bet as a cautionary analogy, while ultimately concluding compute supply is far less elastic than commodity extraction.
Why this matters for AI operators
- Compute pricing is already inflecting upward — plan accordingly. Spot GPU prices are 40%+ above the February 2025 trough, and the Google/SpaceX deal at 2x spot suggests reservation premiums will widen. Operators locking in long-term compute contracts now may be doing so before further price escalation driven by demand that outpaces a structurally capped 3x annual supply growth.
- The wafer-share absorption ceiling is a hard constraint with a known timeline. When AI reaches ~86% of leading-edge wafer capacity (projected end of 2026), the 1.8x supply growth factor from wafer reallocation disappears. Combined with stalling Moore's Law gains and ASML EUV bottlenecks, the 3x annual compute growth may itself become unsustainable — compressing inference capacity relative to demand.
- Model efficiency becomes a first-order economic moat. As compute costs rise, the lab or enterprise deploying the most token-efficient model captures disproportionate margin. For applied AI operators, this argues for aggressive evaluation of model efficiency (output quality per token), not just raw capability benchmarks, when selecting inference providers.
- Low-value AI workloads face a structural pricing squeeze. If frontier labs are willing to pay more per token to run agentic AI research or software engineering automation than commodity content generation, consumer-facing and low-margin AI applications may face rising unit economics — making current cost-based business cases fragile over a 12–24 month horizon.
Full analysis
Dwarkesh Patel ran the arithmetic on frontier compute and landed on an uncomfortable number: if AI hits human-level software engineering, an H100-equivalent GPU should rationally rent for ~$250K a year, roughly 15x today's spot price. His logic is a supply-demand squeeze. Anthropic's revenue has 10x-ed three years running while global compute only grows ~3x a year, and the two ways out (fatter lab margins or higher compute prices) are running out of room on the margin side.
What's actually being decided: whether teams shipping AI into production should treat cheap inference as a temporary subsidy and plan for a world where tokens get materially more expensive. This is a Type 2 call for most reads (you can renegotiate contracts, swap providers, re-architect), but the long-term compute contracts Patel points at are Type 1. Forcing function: the wafer-absorption ceiling he dates to end of 2026. Below, the council pressure-tests the thesis and what an ad-tech engineering lead should do with it.
The Skeptic. Patel's own caveat is the strongest line in the essay: this pattern-matches to every scarcity bet that lost, and he name-checks Simon-Ehrlich to prove he knows it. The tell is the revenue extrapolation. "$1T by end of 2026 if the trend holds" is doing all the work, and three data points of 10x growth is a slope you drew through a startup's early ramp, not a confirmed trend. Revenue 10x-ing does not mean paid demand 10x-es; a lot of that is discounted, subsidized, or free-tier usage that evaporates the moment prices move. For a PM: the scary number assumes AI keeps getting more useful and everyone keeps paying full freight. Break either and the squeeze softens.
The Researcher. The honest number in here is the inference split, not the GPU valuation. OpenAI ran ~25% of compute on inference in 2024, ~50% now, per Epoch AI. That shift is real and measurable, and Patel's read on why labs resist tilting further is the interesting bit: more inference share signals training progress stalled, which undercuts the AGI story they raise money on. So the compute allocation is partly a financing decision, not a physics one. The $250K/GPU figure, by contrast, is a thought experiment resting on "true human-level software engineer on one H100," which no benchmark today supports. SWE-bench gains are real but nowhere near "replace a $250K engineer end to end."
The Open-Source Advocate. Every line of this thesis assumes you're renting frontier tokens from a lab that can outbid you. That's the escape hatch. If compute gets 10x more precious, the pressure to run Qwen, Llama, or a distilled open model on your own reserved capacity goes way up, and for most ad-tech workloads (creative gen, classification, bid-adjacent scoring) an open model at 80% of frontier quality already does the job. Patel's own "Alkan Allen Effect," that the most token-efficient model wins when compute is dear, cuts toward small efficient open weights as much as toward frontier labs. The squeeze he describes is a reason to own your inference, not rent it.
The Compute Pragmatist. The price signals are the part I'd act on. Google paying $900M/month for 110,000 GPUs at 2x spot, and spot itself already 40%+ above the February 2025 trough, tells you reservation premiums are widening right now, independent of whether AGI ever shows up. That's not a forecast, that's a print. The wafer ceiling is the constraint with teeth: 3x annual supply decomposes into ~1.4x Moore's Law, ~1.2x new fabs (gated by ASML EUV output through 2030), and ~1.8x from AI eating smartphone and PC wafer share. That last 1.8x dies when AI hits ~86% of leading-edge wafers, projected end of 2026. After that the 3x itself is in question.
The Builder. What do I do Tuesday? Not much that's dramatic, and that's the point. The consumer/commodity squeeze is the real operational risk: if labs would rather sell tokens to agentic research than to your bulk creative-variation pipeline, the batch jobs where you burn millions of cheap tokens are exactly what gets repriced. So I'd instrument cost-per-outcome per workload today, separate the latency-critical RTB-adjacent calls from the batch generative stuff, and make sure the batch tier can fall back to an open model on owned capacity. For a PM: the AI features whose business case only works at today's token price are the fragile ones. Find them now.
Where the council splits. Three real disagreements. First, the Skeptic and the Compute Pragmatist part ways on why prices rise: the Pragmatist sees a physical supply cap that holds regardless of demand hype, the Skeptic sees a demand curve propped up by subsidized usage that could snap. You can believe prices rise for the first reason and not the second. Second, the Researcher and the whole $250K thesis: the valuation only lands if models actually reach human-level engineering, and the evals don't say that yet. Third, the Open-Source Advocate versus the frontier-lock assumption: Patel's math only bites teams who must rent frontier tokens, and most ad-tech inference doesn't need them.
What it hinges on. Two beliefs. One, does compute supply growth actually stall near the wafer ceiling on Patel's timeline. That one has hard, observable inputs (ASML EUV shipments, TSMC leading-edge allocation) and looks more solid than the demand side. Two, does frontier capability keep justifying frontier prices, or does the open-weight tier close enough of the gap that the squeeze routes around you. For an ad-tech shop, the second is the one you control. The council leans toward "prices go up, but you have an exit": lock reserved capacity for latency-critical work, move batch generative to owned open-model inference, and stop building features whose margin only survives at 2025 token prices.
What to verify before signing anything long: benchmark your top three workloads at output-quality-per-dollar on both a frontier API and a hosted open model, and negotiate any multi-year GPU contract with a repricing or exit clause, because you're contracting into a market Patel himself admits could break either way.
Prediction: Frontier inference list prices (per million tokens on OpenAI's and Anthropic's flagship models) will NOT rise 10x by end of 2026; the headline price of the leading general model will be flat or lower than August 2026 levels when the wafer-ceiling milestone Patel dates to end of 2026 arrives.
Confidence: Medium — competition and efficiency gains have cut per-token prices every year so far.
Why: The thesis needs demand to keep 10x-ing while supply caps out, but the observable trend on published token prices runs the other way: every major model generation since GPT-4 has shipped at a lower per-token price than the one before, driven by distillation, MoE routing, and quantization, not by more wafers. Patel's own "efficient model wins" effect pushes labs to compete on tokens-per-dollar, which shows up to buyers as flat or falling list prices even if raw GPU rental costs rise. The 10x compute-price squeeze can be true at the bare-metal GPU layer while frontier API prices stay flat, because labs eat the gap in margin and efficiency rather than pass a 10x to customers who would defect to open weights. Prices spiking 10x would require the leading lab to have no efficiency lever left and no competitor, which nothing in 2026 supports.
Revisit by 2026-12-31: We're right if OpenAI's and Anthropic's flagship per-million-token prices are at or below August 2026 levels. We're wrong if either lab's flagship list price is 2x or more above August 2026.
Comments