Podcast episode
The Professor of Outputmaxxing — Anjney Midha, AMP
ai-in-adtech cloud-costs engineering
TL;DR
Anjney Midha (CEO of AMP, investor in Anthropic/Mistral/Black Forest Labs) makes a detailed case that the AI race is increasingly a systems-efficiency problem, not a raw GPU procurement problem — with MFU (Model FLOPs Utilization, the fraction of theoretical compute actually used for model training) at best-in-class 60–70% today and reportedly sub-10% at some frontier labs. The episode covers AMP's vision for a compute grid modeled on electric-grid ISO structures, Anthropic's culture and coding breakthrough, and AI applications in end-of-life healthcare prediction. AI infrastructure operators and lab-watchers will find the MFU benchmarks and Anthropic inside-view most actionable.
What was covered
-
Compute utilization crisis at scale. Midha contrasts Google's standard of 96%+ node utilization (where 95% was considered an "outage") and 60–70% best-in-class MFU against what he implies is sub-10% MFU at some frontier labs (xAI cited in the show notes). He attributes the gap to misaligned incentive chains between capital providers and cluster operators, compounded by forced-fast scaling without iterative bring-ups.
-
AMP's "compute grid as electric grid" vision. AMP is building a multi-cloud, multi-silicon compute pooling layer modeled on PJM Interconnect (the U.S. northeast power ISO). Target: 1.2–1.3 GW base-load capacity secured over four years, with a stated need for 6 GW of spike capacity over the same period. AMP's engineering leads (Seb and Mihai) built Google's Borg/GQM scheduler. The grid uses dynamic prioritization ("interruptible demand") to allocate jobs across tenants.
-
Research hoarding at DeepMind as market failure. Midha claims DeepMind has a six-month business-team embargo on papers — if anyone on the commercial side flags a paper as interesting, it is embargoed permanently. He argues this creates negative externalities and an adverse-selection problem (only research deemed commercially valueless gets published).
-
Anthropic's coding breakthrough framed as preparation, not luck. Midha says Anthropic "achieved takeoff" around October of last year (implied: Claude 3.x checkpoint, released December). He attributes the coding lead to four years of extreme resource constraint forcing efficiency, combined with a day-one P0 thesis: crack coding → crack AGI. Anthropic's burn rate was described as dramatically lower than OpenAI's.
-
Data center community backlash risk. Midha cites a claim that up to 20% of planned U.S. data centers this year face community-opposition risk. He proposes passing marginal unit economics (~$0.50/hr above a $4/hr baseline) directly to host communities as a solution.
-
MatX chip co-design strategy. Rainer Pope's MatX adopted NVIDIA's reference rack architecture entirely to avoid fighting on the data center integration front, innovating only on logic-die co-design. Midha frames NVIDIA's open reference architecture as enabling, not threatened by, alternative silicon.
-
AI for end-of-life prediction. Midha's 14-year obsession: using AI/RL on longitudinal patient datasets (Stanford's 12M-patient STRIDE dataset; VA has the only larger one) to give terminal patients precise survival estimates, reducing the 30%+ of Medicare/Medicaid spend on end-of-life care driven by physician malpractice liability uncertainty.
-
Culture as fragile, not moat. Using Anthropic as the case study, Midha argues early-stage hardship and capital scarcity sharply define culture; labs receiving too much capital too early never face the forcing function that crystallizes a P0 and a durable culture.
Notable claims & predictions
-
Anjney Midha: "Best-in-class MFU today is somewhere between 60 and 70%. [xAI running sub-10%] is a leadership question — fundamentally an alignment question between the people funding the cluster and those managing it." (Implies most single-tenant frontier clusters are significantly below best-in-class.)
-
Anjney Midha: "Up to 20% of all data centers this year in the US are at risk [of community backlash preventing bring-up]." (Self-qualified as possibly overstated.)
-
Anjney Midha on DeepMind: "There's a six-month embargo window where if anybody on the business team says 'this could be interesting,' it's embargoed for life. The stuff that gets published is the stuff that's not good enough — it's an adverse selection problem."
-
Anjney Midha on Anthropic: "Anthropic basically achieved takeoff in October of last year. That training run — whatever that checkpoint was — to those of us in the community, especially once post-training was done and released in December, it was clear." (P0 was coding from day one because cracking coding = path to AGI.)
-
Anjney Midha: "We need roughly six gigawatts [of compute spike capacity] over the next four years for all our teams to feel like they were able to keep moving the frontier." (AMP's stated long-term ambition.)
-
Anjney Midha: "Teams who can raise too much money, too fast, too early, who don't have to define their P0 — those cultures end up being the most fragile and brittle. They almost don't even make it to takeoff." (Direct critique of current heavily-funded labs without named targets.)
Why this matters for AI operators
-
MFU as the real scaling lever. The gap between theoretical FLOPs purchased and actual training progress (60–70% best-in-class vs. potentially sub-10% at poorly run clusters) means infrastructure operators who close this gap can deliver frontier-equivalent training at materially lower cost. For anyone procuring or managing GPU clusters, MFU measurement and incentive alignment between capital, ops, and model teams is now a competitive differentiator — not a nice-to-have.
-
Compute market structure is shifting toward ISO/grid models. AMP's framing of itself as an independent system operator — pooling demand from uncorrelated tenants (labs, universities, healthcare AI companies) against contracted base-load supply — is a structural bet that no single frontier lab's internal cluster will be efficient enough to out-compete pooled utilization. If true, this has direct implications for how hyperscalers and neocloud providers price reserved vs. spot capacity.
-
**Alternative silicon (MatX, AMD) adoption path is clearer than assumed.
Full analysis
Decision Council: The Outputmaxxing Episode
Step 1 — Frame
The story: a well-connected investor (boards/checks at Anthropic, Mistral, Black Forest Labs) argues the AI race is now an efficiency problem, not a GPU-count problem. His headline number: best-in-class MFU — the fraction of a GPU's theoretical math throughput actually spent moving a training run forward — sits at 60–70%, while some frontier clusters reportedly run under 10%. He's building AMP, a "compute grid" that pools demand across tenants like a power utility.
What's actually being decided for an AI builder: not "should I care about AMP," but "is my inference and training spend 3–5× more wasteful than it needs to be, and is that the lever I should be pulling instead of buying more capacity?"
Reversibility: Mostly Type 2 for builders. Measuring your own utilization, renting from a different provider, or restructuring a training job are all reversible. Picking a pooled-compute vendor on a multi-year base-load contract is Type 1 — that's the one to slow down on.
Timeline / forcing function: None hard. This is a "re-examine your assumptions" briefing, not a deprecation deadline. The MatX/AMD/silicon angle has a slow clock; the MFU-audit angle you could act on next sprint.
Ad-tech relevance, stated plainly: low-to-moderate and indirect. Nothing here touches CTV, identity, measurement, or the bidstream. The real signal for ad-tech readers is second-order — inference cost curves — and I'll keep that honest rather than inflate it.
Step 2 — The Council
The Skeptic Sub-10% MFU is a great cocktail-party number and almost certainly cherry-picked. MFU isn't one thing — training MFU, inference MFU, and "MFU during a chaotic cluster bring-up" are wildly different regimes, and quoting the worst moment of a competitor's worst week as if it's steady-state is exactly what an investor talking his own book does. AMP's pitch requires incumbents to be wasteful. The PJM-grid analogy is seductive and structurally wrong: electrons are fungible and latency-free; a half-finished training run pinned to specific interconnect topology is neither. For the PM: he's claiming rivals waste 90% of their compute — true at a bad moment, misleading as a headline.
The Researcher The 60–70% best-in-class figure is real and well-documented — published MFU for large transformer runs lands roughly there, and Google's TPU-pod utilization claims are credible. The interesting, checkable claim is the DeepMind six-month publication embargo creating adverse selection: only commercially worthless research gets published. That's a structural argument with teeth — if true, the open literature is being actively skimmed of its best work. The Anthropic "takeoff in October" claim is unfalsifiable vibes. For the PM: the efficiency numbers are solid; the lab-gossip is unverifiable.
The Open-Source Advocate The embargo claim is the most important thing in this episode and nobody's leading with it. If the strongest labs publish only their leftovers, the open-weights ecosystem — Mistral, Qwen, Llama-descendants — isn't 6–9 months behind the frontier, it's behind a deliberately curated frontier. That changes how you read benchmark gaps. The hopeful counter: MatX cloning NVIDIA's reference rack and AMD maturing means the hardware moat is thinning even as the research moat thickens. For the PM: the best AI ideas may now be trade secrets, not papers — so "open" models are racing against a hidden, not a visible, leader.
The Compute Pragmatist This is the lens that matters for ad-tech. If pooled grids and better MFU push effective training cost down, the durable consequence for builders is inference getting cheaper, faster — and that's the only thread that reaches the bidstream. Per-token costs for frontier-class models have already fallen roughly 10× a year; efficiency plumbing like this accelerates it. AMP's actual numbers are sobering: 1.2–1.3 GW base-load with 6 GW of spike demand over four years means the spike-to-base ratio is ~5:1. That's a brutal utility to operate — power grids that bursty are expensive and fragile, which is precisely why his own analogy undercuts him. For the PM: the win for everyone downstream is that running a model keeps getting cheaper; building a "compute utility" to deliver it is much harder than the slide suggests.
The Builder The episode I can act on is buried in the Granola notes, not the podcast. "Workflows start at 20–30k tokens vs. 500k" and "agents burn tokens fast" — that's outputmaxxing at my altitude. The lesson isn't AMP's gigawatts; it's that MFU has a personal-scale analog: most of my token spend is wasted context I never needed. Before anyone romanticizes grid economics, audit your own agent loop — that's the 5× sitting in your codebase today. For the PM: the same "you're wasting 90% of your compute" problem exists inside your own AI features, and you can fix that this week without a utility company.
Step 3 — The Tensions
Skeptic vs. Researcher — is the headline number a measurement or a weapon? The 60–70% best-in-class figure survives scrutiny; the sub-10% rival figure is an investor selecting the worst regime of a competitor and quoting it as steady-state. Same metric, opposite epistemic status. A reader who takes both at face value will over-update.
Open-Source Advocate vs. Compute Pragmatist — which moat is dissolving? The Advocate says hardware is opening up (MatX, AMD) while research is being locked behind embargoes — so the gap shifts from chips to secrets. The Pragmatist says it doesn't matter, because falling inference cost commoditizes capability regardless of who publishes. One sees a widening knowledge gap; the other sees a closing price gap.
Builder vs. everyone — wrong altitude. The entire episode operates at the gigawatt scale, but the only actionable efficiency lever for 99% of the audience is token discipline in their own agent loops. The grandeur of the grid story distracts from the boring win sitting in your own logs.
Step 4 — Synthesis
For an ad-tech / publisher / agency audience, be blunt: direct impact is low. No identity, measurement, CTV, or privacy content here. Don't restructure anything on the strength of this episode.
The decision hinges on one belief that matters and one that doesn't:
- Matters: Does AI efficiency keep compounding inference cost down ~10× a year? The council leans yes — the 60–70% MFU ceiling and the multi-cloud pooling pressure both point at continued downward cost pressure. For ad-tech, that's the only load-bearing takeaway: real-time creative generation, bid-time LLM reasoning, and conversational ad surfaces keep getting economically viable, just on the existing trend line. This episode is confirmation, not news.
- Doesn't matter (yet): Will AMP's grid model win? Type 1 decision, no forcing function, and the Skeptic + Pragmatist both flagged the 5:1 spike-to-base ratio as the analogy eating itself. Watch it; don't bet on it.
What to actually do: ignore the gigawatts, audit your own MFU-equivalent — token waste in agent and RAG (retrieval-augmented generation) loops. Your own meeting notes already found 20–30k vs. 500k-token workflows; that 15× is the real, reversible win, and it's in your codebase, not a power substation.
The one claim worth tracking independently is the DeepMind publication embargo. If frontier labs are systematically withholding their best research, every benchmark-gap estimate between open and closed models is wrong — and that does eventually reach ad-tech buyers choosing between a hosted frontier model and a cheaper open-weights one.
Step 5 — The Prediction
Prediction: By the time the next major Gemini model ships (Google's expected late-2026 frontier release), Google DeepMind will not have publicly confirmed a blanket six-month commercial-veto embargo on research publication, and no on-the-record DeepMind source will corroborate Midha's "embargoed for life" characterization.
Confidence: Medium — Single-sourced investor claim about a rival; labs neither confirm nor deny such policies.
Revisit by 2026-12-31: We're right if no DeepMind employee or official statement corroborates a permanent commercial-veto embargo on the record. We're wrong if Google confirms such a policy, or a credible insider account (current/former staff, on record) substantiates the "embargoed for life" framing.
The claim is structurally plausible — labs do delay sensitive research — but the specific "interesting-equals-banned-forever" version is the kind of vivid detail that gets sharper in the retelling. Labs almost never confirm internal publication policy either way, which is exactly why the falsifiable bet is on the absence of corroboration. For ad-tech readers, this matters only insofar as it tells you whether the open-vs-closed model gap you're pricing into vendor decisions is bigger than the leaderboards
Comments