Refacto AI

Industry story

Inference Revenue Reaches $100B Per GW Per Year, Justifying Massive Power Spend

cloud-costs gpu-supply inference model-pricing

SemiAnalysis argues that the economics of AI inference — generating outputs from a trained model in real time — now make almost any power cost justifiable for frontier labs. The firm estimates inference API revenue can yield $100 billion per gigawatt of datacenter capacity per year at 90%+ gross margins, meaning a $5 billion power plant pays for itself in roughly 20 days of inference revenue. This economic calculus explains why labs like Anthropic and peers are willing to pay 2x more or accept 30% lower efficiency to accelerate deployment timelines.

Analysis

Showing the shorter version.

Your draft

SemiAnalysis put a number on why the AI labs treat power costs like a rounding error: a gigawatt of inference capacity generates $100 billion a year at 90%-plus gross margins. At those numbers, a $5 billion power plant pays back in about 20 days of API revenue. If that's even roughly right, it explains why Anthropic and its peers will pay a 2x premium or accept 30% worse efficiency just to deploy faster.

The 20-day payback figure is almost certainly fiction, and the labs know it. Real inference clusters run at 40% to 60% utilization; push higher and tail latency blows up API contracts. Depreciation on GPU hardware that goes obsolete in three years doesn't show up in the "90% margin" framing either. GPT-4-class pricing fell roughly 95% in 18 months, and three well-funded labs are actively cutting prices to buy ecosystem share. You cannot hold "current API rates" and "competitive market" in the same model.

But the buildout is rational anyway. Power procurement carries a two-year lead time that no chip efficiency gain shortens. Whoever controls gigawatts in 2027 controls who can serve frontier models at scale. The land-grab makes sense even at a near-term loss, because the alternative is showing up to a capacity auction with no chips in hand.

For enterprise buyers, the $100B/GW margin figure is leverage, not a threat. That much headroom exists to be competed away. Don't lock long-term at today's rates. Push for price-step-down clauses and qualify a second vendor before you need one.

The same competitive pressure that cuts your bill also compresses the evaluation cycles before a model reaches you. Cheaper and less-tested arrive together.

The call: Blended per-token prices for frontier-tier models (GPT-5-class, Claude Opus-class, Gemini Ultra-class) fall at least 40% from September 2026 levels by end of Q3 2027. Confidence is medium. Three well-capitalized labs are subsidizing inference to win developer lock-in, and the fat margin in the SemiAnalysis figures is precisely the headroom a competitor undercuts to still profit. The only thing that stops it is a genuine capacity crunch severe enough that labs start rationing access, and that's the less likely path inside a 12-month window. Model your API costs on continued price declines. Get the second provider qualified now.

Also covered this issue

Comments