Refacto AI

Podcast episode

Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028

gpu-supply inference model-pricing

Dylan Patel of SemiAnalysis joins Dwarkesh Patel to argue that Anthropic and OpenAI will control the majority of the world's AI compute by 2028, essentially cornering the market on the infrastructure that runs modern AI.

The structural claim is that frontier labs convert compute into revenue far more efficiently than anyone else: Patel estimates Anthropic generates around $50 million per megawatt today versus roughly $12 million for a cloud provider renting bare metal. That spread lets the labs outbid everyone else for every new data center watt. The inference Patel draws from a flattening in Anthropic's ARR growth while compute spending keeps climbing is that they're quietly reallocating capacity from selling API access to running internal training runs. It's a reasonable inference. It's also one of three or four things that gap could mean.

The 20x supply build Patel's own model projects is the thesis killer. Scarcity funds the corner; abundance breaks it. Watch your own 429 error rates and API latency week over week. That's the first place the squeeze shows up, before any announcement.

Full analysis

The Skeptic. The revenue-per-megawatt story is the single structural assumption the whole analysis rests on, and it's an unaudited number from a research shop selling access to its models of these companies. $50M/MW for Anthropic today, $70-80M by end of 2027, $100M+ after. Nobody outside the labs can check the denominator. And notice the sleight of hand: ARR growth "plateaued" while compute climbs, so the marginal watts must be going to R&D. Or Anthropic hit a demand ceiling, or margins compressed, or they overbuilt. "Compute up, revenue flat, therefore secret training run" is one story that fits the data. For a PM: this is a smart guy inferring a secret from an accounting gap, and the gap has three other explanations.

The Compute Pragmatist. Strip the drama and one thing holds: whoever converts a watt into the most dollars wins the auction for that watt. That part is just arithmetic. If a lab clears $50M/MW selling tokens and a neocloud clears $12M/MW renting bare metal, the lab outbids the neocloud every time compute is scarce. So the direction is right even if the numbers are soft. What breaks the thesis is supply, not demand. Vera Rubin, more fabs, more power coming online. Patel's own model has world compute going from ~10 GW to ~200 GW. A 20x supply build is not a market where two buyers quietly corner everything. Scarcity funds the outbidding; abundance kills it.

The Researcher. The genuinely useful, checkable claim here is the inference-to-training reallocation, and it's testable without any of Patel's spreadsheets. If Anthropic and OpenAI are pulling watts off serving to feed internal R&D, you'd see it in your own logs: rising 429s, longer tail latency, tighter rate limits on the top-tier models, quotas that don't loosen when you ask. The China numbers are also concrete and the most defensible in the episode. China at under 10% of incremental compute, quality-adjusted maybe 20 GW of US-equivalent by 2029. That gap is real and it compounds. Everything downstream of "revenue per MW" is a forecast; the export-control bifurcation is a fact.

The Open-Source Advocate. This is the episode's blind spot, and it's a big one. Both Patels conclude centralization is "irreversible" and admit "no decentralized counter-scenario is identified." That's not analysis, that's a failure of imagination. The whole thesis assumes you need frontier compute to get useful work done. For most production workloads you don't. Qwen, Llama, Mistral, DeepSeek variants already clear the bar for classification, extraction, routing, summarization, most RAG. If frontier API inference tightens and repriced upward, the rational move is exactly the one Patel treats as impossible: pull your commodity workloads off the frontier and run open weights on cheaper, more available silicon. Scarcity at the top is the best recruiting poster open models ever had.

The Builder. Forget 2028. What do I do this quarter. First, instrument for the squeeze now, because it costs almost nothing: track your 429 rate, p99 latency, and effective throughput per dollar on every frontier model you call, week over week. If Patel's reallocation is real, you'll see it before any press release. Second, tier your workloads. Anything that doesn't need the frontier gets an escape hatch to an open model behind the same interface, so switching is a config change, not a rewrite. Third, do not sign a multi-year compute or pricing commitment on the assumption that today's rate holds. That's the one genuinely irreversible mistake in reach of a normal team.


Where they split:

The Compute Pragmatist and the Skeptic both think the 20x supply build undercuts the corner-the-market story, but Patel's rebuttal lives in the same episode: the labs absorb 40-50% of incremental compute, so supply growing doesn't help you if they're eating half of every new gigawatt. That tension is the real fork. Does new supply relieve the market, or do the labs just eat it as it arrives.

The Researcher and the Open-Source Advocate agree the inference squeeze is the checkable claim, but split on what it means. The Researcher watches it as a warning about frontier availability. The Advocate reads the same squeeze as the trigger that finally pushes production workloads off the frontier and onto open weights. Same signal, opposite response.

What it hinges on: one belief, and it's testable from your own dashboard. Are the frontier labs actually reducing the share of compute they sell as inference? If yes, API-dependent products get squeezed on price and availability regardless of whether the sovereign-debt story ever materializes. If no, the whole centralization thesis is a forecast dressed as a trend. You don't need Patel's numbers to answer this. Your rate-limit logs answer it.

Which way the council leans: the macro apocalypse is unfalsifiable coffee talk, and the "irreversible two-lab world" ignores that most work doesn't need frontier compute. But the near-term inference squeeze is plausible, cheap to hedge, and already checkable. Instrument for it, build the open-model escape hatch, don't sign long.


Prediction: Between now and the GPT-6 and next-Claude frontier releases expected by mid-2027, at least one of OpenAI or Anthropic will impose a materially tighter constraint on paid top-tier API access, a hard rate-limit cut, a new usage tier that gates the best model, or a per-token price increase on the flagship, that developers will publicly complain about as a capacity squeeze rather than a routine pricing change.

Confidence: Medium. The incentive is real and one-directional, but timing is theirs to control.

Why: Patel's checkable claim is that frontier labs earn far more converting a watt into training progress than into sold tokens, and that Anthropic already began shifting compute off inference in the last three months. If that economics holds, the labs face a standing incentive to ration external inference on the flagship whenever their own R&D wants the watts, and rationing shows up to developers as rate cuts, gated tiers, or price hikes on the best model. We have already seen the pattern, capacity-driven rate limits and "priority" tiers appear whenever demand outruns supply. The opposite outcome, both labs keeping flagship access cheap and unconstrained through two release cycles, requires them to leave the higher-value use of their scarcest resource on the table, which is the less likely behavior for companies this compute-bound.

Revisit by 2027-06-30: We're right if OpenAI or Anthropic ships a flagship rate-limit cut, a new gate on top-tier model access, or a flagship per-token price increase that draws public developer complaints about capacity before then. We're wrong if both keep flagship API access at current-or-looser limits and current-or-lower flagship token prices through that window.

One more thing for your team: the hedge here is asymmetric. Instrumenting your 429 rate and building an open-model fallback behind your existing interface costs a sprint. Getting caught mid-rewrite when your flagship provider gates you costs a quarter. Cheap insurance against a call that's only Medium confidence is still worth buying.

Comments