Refacto AI

Podcast episode

How Big Is the AI Economy?

The AI economy is real — $175B annualized, token volumes up 14× — but the unit economics underneath it are quietly shifting on everyone. Token prices collapsed from $17 to ~$2 per million on the commodity tier, but frontier model pricing has stopped falling, AWS contract blocks are up 20%, and memory costs just spiked 60%. Amazon already lost its wholesale Claude deal and is scrambling back to its own Nova models. Any ad-tech shop that built margin assumptions on today's inference prices should check those assumptions now.

Full analysis

The Exponential View report says the AI economy is a $175B annualized business growing 3× faster than any prior IT wave, with token prices collapsing from $17 to a projected $2 per million while capability keeps climbing. Buried in the same episode: Anthropic clawing back Amazon's wholesale compute deal, Meta banning rival coding agents to dodge distillation lawsuits, AWS hiking GPU rental 20%, and Zuckerberg quietly telling staff agents haven't panned out. That's the tension worth unpacking.

Reversibility: Mostly Type 2 for builders — you can reprice, reprovision, and re-audit workflows on a quarter's notice. The one Type 1 lurking here is committed GPU capacity contracts, which lock you in for a year-plus.

What's actually being decided: Not "is AI real" — the revenue data settles that. The live question is whether the unit economics you built your product on survive the next repricing cycle, and whether the "wholesale sweetheart deal" era ending changes your cost model.


The Skeptic — The report is a beautiful narrative and I don't trust round narratives. "$1B in cumulative revenue every 180 days now takes under 2 days — a 90× acceleration" is the kind of stat you cite when you're selling a thesis, not testing one. Notice what's stapled to the same episode: Zuckerberg told his own staff agents "haven't progressed the way" they hoped, and Meta's own Applied AI division is banned from using the best coding agents. The company with the most compute and the most incentive to believe is privately hedging. The 14× token-volume growth is real, but "agentic coding uses 1,200× the tokens of a chat task" means volume growth is partly an artifact of wasteful loops, not value creation. For the PM: booming usage numbers can mean the product got great, or that it got chatty — those aren't the same thing.

The Researcher — The single most important number for ad-tech is the price-capability scissors: blended price per million tokens falling $17 → $2 (an 8.5× drop in two years) while the capability index climbs 112 → 158. That's roughly a 12× improvement in intelligence-per-dollar. For anyone doing creative generation, contextual classification, or brand-safety scoring at scale, that curve is your roadmap — jobs that were uneconomic at $17 become trivial at $2. But read the Anthropic quote carefully: "we reduced Opus pricing in November 2025 and that price has held." Held, not fallen. The cheap blended average hides a frontier tier whose price is now sticky. For the PM: the AI that's getting cheaper fast is the good-enough tier; the very best model stopped getting cheaper.

The Compute Pragmatist — This is the story that actually moves ad-tech P&Ls. Spot H100 prices down 40% from the May peak, but AWS contract blocks up 20% and Semi Analysis confirms term pricing is rising. That's not softening demand — it's demand graduating from experiment to production and locking in capacity. If you're running programmatic bidding models or real-time creative gen on on-demand pricing, you're on the wrong side of that shift: cheap when idle, brutal when you scale. And "Ramageddon" — Micron memory up 60% in three months, 4× YoY — hits everyone with a KV-cache and a context window, because long context is a memory problem. Apple raising prices 15% and begging to buy blacklisted Chinese memory tells you the squeeze is real, not a spot blip. For the PM: the chip shortage moved from GPUs to the memory that feeds them, and that raises the floor on every inference bill.

The Open-Source Advocate — Watch the Amazon move: losing its wholesale Claude deal, it's dusting off in-house Nova and shopping OpenAI as a hedge. That's the whole ad-tech playbook in miniature. When your frontier vendor normalizes you to market-rate token pricing, the counterweight is a good-enough open or in-house model for the 80% of tasks that don't need frontier reasoning — classification, tagging, summarization, first-draft creative. Meta banning Claude Code isn't just legal caution; it's a tell that the labs now treat their outputs as a moat worth litigating over, which raises the strategic value of weights you actually own. For an ad-tech shop, the question isn't "Claude or GPT" — it's "which tier of work can I move to Llama/Qwen/Mistral before the next renegotiation." For the PM: owning a mediocre model you control beats renting a great one whose price your vendor sets.

The Builder — The meeting notes here are the real world the report abstracts away. Claude's now wired into Notion, GitLab, Jira, and Slack — but the team's stuck on a Teams-vs-Enterprise account tier that blocks DLP and DMs, and a "configure button to make Claude more adversarial when it's too compliant on reviews." That's the actual state of enterprise AI: capability is abundant, plumbing and permissions are the bottleneck. On the ad-tech-specific "QuantumPath" positioning: buyers are paralyzed — "Should I do Scope3? Newton Research? Something else?" That paralysis is the market reality that the $175B number papers over. Demand is real; buyer confidence in what to buy is not. Warner's agent-neutrality bill, if it ever moves, would force platforms to accept third-party agents with a "duty of loyalty" — that reshapes who controls the shopping/bidding interface, which is existential for anyone building agentic bidding. For the PM: the tech works; the org chart, the license tier, and the buyer's fear are what's slowing you down.


Where they split:

  1. Is the growth curve demand or plumbing? The Researcher and Compute Pragmatist read the token-price scissors as a genuine, durable productivity unlock. The Skeptic reads a chunk of the 14× token growth as agentic loops burning tokens and Zuckerberg's own hedging as the tell. Both can be true: the economy is real and a fraction of the usage is waste that reprices when someone checks the bill.

  2. Does cheap inference save you or trap you? The price-per-token collapse says margins improve. The Compute Pragmatist says the frontier tier's price has held, contract GPU and memory costs are rising, and the savings only accrue if your workload sits on the commoditizing tier — not if you need the best model in real time.

  3. Who owns the agent interface? The Open-Source Advocate and Builder see Amazon's Nova hedge and Warner's neutrality bill converging on the same point — control of the model and the agent surface is the strategic asset. If duty-of-loyalty rules land, the agent works for the user, not the platform, which upends the take-rate logic of every intermediary in the ad chain.

What it hinges on: For ad-tech specifically, the decision hinges on two facts you can actually check. First, which tier your workload needs — if brand-safety scoring and creative gen run fine on the commoditizing $2 tier, the price curve is pure tailwind; if you need frontier reasoning in the bidding loop, you're exposed to sticky frontier pricing plus rising compute. Second, whether you're on committed or spot capacity — the market just told you committed is where production is going, and spot is a trap at scale.

Which way the council leans: The AI economy is not a bubble in the "no revenue" sense — that argument is dead. But the era of subsidized pricing is ending, and the cost structure is quietly shifting from cheap-experiments to expensive-production. Ad-tech's near-term impact is moderate and mostly on the cost side, not a demand revolution. The bigger latent threat is regulatory: agent-neutrality and duty-of-loyalty rules, if they move, hit intermediary economics harder than any pricing change.

What to verify before committing: Audit your token spend by task tier this quarter — separate what genuinely needs frontier from what's running on it out of laziness. Model your inference bill against a scenario where your vendor moves you to market-rate token pricing (Amazon just got that call). And if you're building or training internal models while using Claude/Codex operationally, audit for distillation exposure now — Meta didn't ban those tools for fun.


Prediction: By the time Q4 2026 hyperscaler earnings are reported (late Jan / early Feb 2027), at least one major cloud or hyperscaler besides Amazon will publicly confirm it is shifting frontier-model workloads toward an in-house or cheaper-tier model to hedge against rising per-token costs.

Confidence: Medium — The Amazon-Anthropic repricing plus rising contract GPU and memory prices make cost-hedging a near-universal incentive.

Why: Wholesale sweetheart deals are ending industry-wide, AWS already hiked GPU rental 20%, and memory costs ("Ramageddon") are climbing 4× YoY — every large buyer now has the same math problem Amazon does, and in-house/cheaper-tier hedging is the obvious response labs have already telegraphed.

Revisit by 2027-02-15: We're right if a second hyperscaler or major AI buyer (Microsoft, Google, Meta, or a top-tier enterprise) publicly confirms shifting frontier workloads to in-house or cheaper models to control cost. We're wrong if frontier per-token pricing stays flat or falls and no such hedge is announced

Comments