Podcast episode
What the Best Business AI Users Are Doing Different
agents cost-compression inference model-pricing orchestration
Nathaniel Whittemore's AI Daily Brief covers where enterprise AI budgets are actually moving, using a KPMG survey of 2,100 executives and spend data from RAMP. The short version: open-source models (Llama, Qwen, DeepSeek and their cousins) account for under 5% of business AI spend, and the price war between OpenAI and Anthropic is driving that number, not open weights.
The finding worth acting on is model routing. Only 23% of organizations can automatically send a query to the cheapest model capable of handling it. That's the capability that turns falling token prices into savings you actually keep. The firms without it pay whatever their one hardcoded model costs, even as prices drop around them.
The management layer above the model is where the edge is being built right now. Model selection is mostly commoditized. Before your next renewal, find out whether your stack can route. If it can't, that's the gap to close.
Full analysis
Two things in this episode tell you where enterprise AI money is actually going, and they point in the same direction. Token prices are falling even as usage climbs. And the companies getting real returns aren't winning on model choice. They're winning on the plumbing above the model: routing queries to the right model, governing what agents can do, keeping their data where they want it. The rest of the episode is spectacle. Space chips, trillion-dollar IPOs, Trump asking Grok about Venezuela. Fun, but not where you should spend your planning time.
What's being decided, for a reader who buys and builds with AI: is the edge in picking the best model, or in building the layer that manages models? The KPMG numbers say the second. That's a change in where to put your next dollar.
The Skeptic
RAMP's Eric Karazian says enterprise spend fell 5.2% in a single week while token volume rose, and pins it on OpenAI versus Anthropic price-cutting. One week is noise, not a trend. Watch whether it holds for a quarter before you rebuild your budget around it.
The KPMG survey is self-reported by 2,100 executives grading their own maturity. "We have a formal AI management layer" is exactly the kind of thing a senior leader says yes to whether or not it's true. The 86%-versus-31% gap is real as a vibe, soft as a fact.
And host Nathaniel Whittemore waving off Anthropic's $8 billion loss because 2026 revenue might hit $46 billion? That's Whittemore talking his book. A 10x projection is a hope, not a line item.
The Researcher
The one number worth building on is model routing: 23% of all organizations can send a query to the cheapest model that can handle it. That's the capability that turns a price war into savings you actually keep. If OpenAI and Anthropic keep cutting, the firms with routing capture the drop automatically. The other 77% pay whatever their one hardcoded model costs.
Claude 5.5 sightings matter less than people think. Whittemore himself notes that since Opus 5.5 is already strong, the jump will feel small even if the underlying gains are real. We've hit the stretch where each new top model moves the benchmark but not your Tuesday workflow.
The agent-strategy shift is the genuine finding. Revenue-focused agent use is rising, pure cost-cutting is falling, and most firms now do both.
The Open-Source Advocate
Here's the line that should stop you: open-source models are under 5% of business spend, and Eric Karazian says the price war is "driven almost exclusively" by the two closed labs. So much for the year of open weights in the enterprise. Qwen, Llama, Mistral, DeepSeek-V4.1-Flash, whatever you like, they're not moving enterprise budgets yet.
But read OpenAI's Dev Day move carefully. They're letting enterprises spend their OpenAI commitment on third-party open-weight models bought through OpenAI's marketplace. That's open models arriving through a closed front door. You get variety; OpenAI keeps your wallet. It props up open-source usage while making sure the money still routes through one platform. Convenient for OpenAI. A narrower door than running the weights yourself.
The Compute Pragmatist
Falling cost-per-token is now structural, not a promo. Two labs with billions in losses are fighting for share, and that fight flows straight into your bill. Model your costs assuming next year's token is cheaper than this year's. Plan the savings; don't bank them until you can route to capture them.
Google's TPUs in orbit are a decade-out science project. Four chips, 15-minute power bursts, a soccer-field satellite that needs Starship to exist. Google's own Travis Beals says the win condition is that nobody notices. Nothing here changes your inference bill before 2030. File it under "someday," not "roadmap."
The real near-term compute story is boring and correct: the management layer is where you control spend, and the chip is beside the point.
Where they disagree
The Skeptic and the Researcher split on the KPMG survey. The Skeptic says self-graded maturity is soft. The Researcher says routing at 23% is a hard, checkable capability regardless of what else executives claim. They're both right, and that resolves the decision: trust the capability numbers (routing, agent adoption), discount the governance-vibe numbers.
The Open-Source Advocate and the Compute Pragmatist split on OpenAI's marketplace. Is it open models winning, or OpenAI locking you in? Both. You get flexibility at the model layer and dependence at the billing layer. Whether that's a good trade depends on how much you'd pay in engineering time to run open weights yourself.
What it hinges on
One belief: does falling token price actually reach your budget? It only does if you can route. The firms with routing turn the price war into margin. The firms without it watch cheaper tokens exist somewhere else. Before your next renewal, run one test: take your three highest-volume workflows and check whether a cheaper model clears your quality bar on them. That's the whole game this quarter. Not which frontier model tops the leaderboard.
Prediction: Enterprise AI spend on open-weight models, as measured by RAMP's token-spend data, will still be under 10% of total business AI spend when RAMP reports its Q1 2027 figures in April 2027.
Confidence: Medium — open-source is starting under 5% and the two closed labs own the price war.
Why: RAMP's economist Eric Karazian says open-source is under 5% of business spend today and that falling prices are driven "almost exclusively" by OpenAI and Anthropic competing, not by open models undercutting them. When the closed labs are the ones cutting prices, the usual reason to switch to open weights, cost, weakens, because the closed option is getting cheaper on its own. OpenAI's new marketplace makes this worse for the open-weight share by letting enterprises consume third-party models while keeping the spend counted inside OpenAI's ecosystem, so even "open" usage routes through a closed platform's billing. For open-weight share to double to 10%+ in two quarters, you'd need either a cost crisis the price war is actively preventing or a capability jump that closed models keep matching.
Revisit by 2027-04-30: We're right if RAMP's Q1 2027 data shows open-weight models under 10% of business AI spend. We're wrong if they hit 10% or more.
Comments