Podcast episode
The Most Important Trends Showing Up in New AI Products
agents cost-compression inference model-pricing orchestration
"The Most Important Trends Showing Up in New AI Products" is a roundup from Nathaniel Whittemore and Dillon Rolnick covering the week's AI product moves. The headline: cheap frontier-quality models are now American, and even the big closed labs are routing users to whichever model wins the task rather than defending their own.
The specific numbers matter here. Anthropic's Haiku 5.5 (a small, fast, cheap model) hits 72.4% on OSWorld, a benchmark for whether a model can drive a computer by clicking and typing, against 48.9% for GPT-6 Luna. Real-task cost runs 12 to 21 cents. Rolnick's Nous Research, an open-source model shop, is tracking to $100M revenue and just hit a $1.5B valuation. Meanwhile, SpaceX is shopping a two-page deck to raise $40B for NVIDIA chips, and its credit default swaps price a 15% default chance by 2031.
The cheap-inference story is real, but the cost floor is being held up by leveraged financing. Those two things are in conflict.
Analysis
Showing the shorter version.
Raw Model Capability Is No Longer the Competition
Two things happened this week that matter more than the billion-dollar headlines. The cheapest frontier-quality models are now American. And even the labs with the best models are routing you to whoever wins the task. Both point the same direction: the model underneath is a commodity. The price and the product are what matter.
The Anthropic Haiku 5.5 numbers are specific. 72.4% on OSWorld, a benchmark for whether a model can drive a computer by clicking and typing, against 48.9% for GPT-6 Luna. More than double GPT-6 Luna on TerminalBench 4.0. Those are real gaps on tasks operators actually run. The pricing is the louder signal: 25% of Sonnet's per-token cost, with an 80% discount on prompts under 100K tokens. Artificial Analysis pegs real task cost at 12 to 21 cents per task. Haseeb Qureshi's framing, "buy American for the cheapest LLMs," is a genuine reversal. DeepSeek's price edge is gone.
The routing story confirms what the pricing implies. xAI built a frontier model and still routes Grokbot to Claude and Midjourney. Elon Musk's own product team is conceding no single model wins everything. The RAMP data, which tracks what companies actually pay for AI each month, shows OpenRouter and Featherless, both model routers, showing up in enterprise spend. Operators are already voting for optionality with budget. The right builder response is to stop writing code against one vendor's SDK and put a router in front, so swapping Haiku for whatever beats it next month is a config change.
Nous Research hitting a $1.5B valuation reinforces this. 22 million installs, roughly 2.5% of all tokens used globally, tracking to $100M revenue by year-end. CEO Dillon Rolnick's pitch is observability and no lock-in. That pitch is landing precisely because the routing trend proves his point.
OpenAI's Intelligent UI, charts and interactive output rendered inline, gets attention. It's trivially copyable. Anthropic or Google ships the same thing in a quarter. If you've built your own front end, it does nothing for you. If you ship inside ChatGPT, it just raised the bar on what users expect. Not nothing, but not a moat.
The part the cheap-inference story hides is the financing underneath it. SpaceX is reportedly shopping a two-page deck to raise $40B for NVIDIA chips, and its credit default swaps are pricing a 15% chance of default by 2031. Oracle is seeking private credit under downgrade risk. Broadcom is lending its own credit rating to OpenAI's roughly $50B custom-chip build. Haiku at 12 cents a task looks like deflation, but the hardware underneath is financed like a leveraged buyout. If any of these deals wobble, spot inference prices do not stay this low.
That is the real disagreement. The cheap-inference camp sees a genuine price collapse: American labs closed the gap, open models have a business, routers let you always buy the cheapest winner. The skeptical camp sees low prices propped up by junk-grade debt and says the floor is borrowed. Both cannot be right for long.
The call: At least one of the three major AI chip financing deals (SpaceX's $40B NVIDIA raise, Oracle's private-credit chip purchase, or Broadcom's credit-backed OpenAI silicon deal) gets downsized, delayed past its reported timeline, or repriced at materially worse terms before Q1 2027 earnings land in February 2027. Medium confidence. The debt is already being priced at distress levels. The opposite outcome requires private-credit appetite for AI infrastructure to hold at a peak through a quarter when the CDS market is already doing the doubting.
The practical question is whether 12-to-21-cent frontier inference is a stable fact to build on or a promotional price held up by debt. Run the test. Rebuild one real workload on Haiku 5.5, measure the task cost yourself, and put a router in front so you are not married to it. If the price is real, you win. If it's borrowed, you're hedged.
Two things happened this week that matter more than the parade of billion-dollar headlines. First, the cheapest frontier-quality models are now American. Second, even the labs with the best models are giving up on model loyalty and routing you to whoever wins the task. Both point the same direction: raw model capability is no longer where anyone competes. The price and the product are.
Here's the frame. The decision facing anyone who buys or builds with AI is whether to keep betting on a single model vendor or to design for a world where the model underneath your product swaps out monthly. That's an easy choice to undo if you build for it now, and an expensive one to undo if you've hard-wired one vendor into your stack. Nothing sets a hard deadline here, but the price moves are happening on a monthly cadence, and the RAMP enterprise spend data (which tracks what companies actually pay for AI each month) shows share moving that fast too.
The Skeptic
NLW's claim that "product differentiation is going to really matter" because intelligence is now "on tap" is the kind of line that sounds true and ages badly. We heard the same about cloud in 2012. Differentiation didn't vanish. It moved to the companies that owned distribution. OpenAI's Intelligent UI (charts and interactive bits rendered inline instead of walls of text) is nice. It is also trivially copyable. Anthropic or Google ship the same thing in a quarter. And Jamie Dimon saying cyber risk "went up tenfold after Mythos" is a bank CEO who sells nothing by understating threats. Treat the number as marketing, not measurement.
The Researcher
Look at what Anthropic actually published on Haiku 5.5, because the numbers are specific. 72.4% on OSWorld, a test of whether a model can drive a computer by clicking and typing, against 48.9% for GPT-6 Luna. More than double GPT-6 Luna on TerminalBench 4.0, a coding-agent test. Those are real gaps on tasks operators actually run, not trivia-quiz benchmarks. The pricing is the louder signal: 25% of Sonnet's per-token cost, with an 80% discount on prompts under 100K tokens. Artificial Analysis pegs real task cost at 12 to 21 cents. Haseeb Qureshi's "buy American for the cheapest LLMs" is a genuine reversal. DeepSeek's price edge is gone.
The Open-Source Advocate
Don't let the American-labs-won narrative bury the actual open-source story, which is Nous Research hitting a $1.5B valuation. 22 million installs, roughly 2.5% of all tokens used globally, tracking to $100M revenue by year-end. CEO Dillon Rolnick's pitch is observability and no lock-in, and that pitch is landing because the routing trend proves his point. When xAI's own Grokbot routes to Claude and Midjourney, Elon Musk is admitting no single model wins everything. That is the open-source argument made by a closed lab. The RAMP data shows OpenRouter and Featherless, both model routers, in the spend charts. Operators are already voting for optionality with budget.
The Compute Pragmatist
Follow the financing, because it tells you the cost floor is being held up with borrowed money. SpaceX is reportedly shopping a two-page deck with pictures of space to raise $40B for NVIDIA chips, and its credit default swaps now price a 15% chance of default by 2031. Oracle is seeking private credit under downgrade risk. Broadcom is lending its own credit rating to OpenAI's $50B custom-chip build. This is the part the cheap-inference story hides. Haiku at 12 cents a task looks like deflation, but the hardware underneath is being financed like a leveraged buyout. If any of these deals wobble, the spot price of inference does not stay this low.
The Builder
What would I actually ship Tuesday? Haiku 5.5 for anything high-volume and latency-sensitive: classification, extraction, computer-use agents that click through web apps. The 80% discount under 100K tokens rewards you for keeping prompts tight, which you should be doing anyway. The routing trend means I stop writing code against one vendor's SDK and put a router in front, so swapping Haiku for whatever beats it next month is a config change, not a rewrite. OpenAI's Intelligent UI matters less than it looks. If you've built your own front end, a lab generating charts inline does nothing for you. If you ship inside ChatGPT, it just raised the bar on what users expect.
Where the council splits
The real disagreement is whether cheap inference is durable or borrowed. The Researcher and the Open-Source Advocate see a genuine price collapse: American labs closed the gap, Nous proves open models have a business, routers let you always buy the cheapest winner. The Compute Pragmatist sees the same low prices propped up by SpaceX junk-grade debt and Oracle credit risk, and says the floor is not real. Both can't be right for long.
The second split is Lambert versus Dimon on Mythos. Nathan Lambert argues that if Mythos-class capability (Anthropic's most capable cyber model, now opened to security professionals in tiers) leaked as open weights, the world would be "more or less fine," an acceleration not a step change. Jamie Dimon says risk went up tenfold. One of them is pricing the danger correctly, and the tiered-access program with mandatory data retention only on the least-guarded tier tells you Anthropic is hedging toward Dimon.
What it hinges on
The question for operators is simple: is 12-to-21-cent frontier inference a stable fact to build on, or a promotional price held up by debt that reprices when the financing tightens? Run the test. Rebuild one real workload on Haiku 5.5, measure the task cost yourself against Artificial Analysis's 12 to 21 cents, and put a router in front so you are not married to it. If the price is real, you win. If it's borrowed, you're hedged.
Prediction: At least one of the three major AI chip financing deals flagged this week (SpaceX's $40B NVIDIA raise, Oracle's private-credit chip purchase, or Broadcom's roughly $50B credit-backed OpenAI silicon deal) will be downsized, delayed past its reported timeline, or repriced at materially worse terms by the time Q1 2027 earnings land in February 2027.
Confidence: Medium. The debt is being priced at distress levels already.
Why: SpaceX's credit default swaps are reportedly pricing a 15% default probability by 2031 and the raise is being pitched on a two-page deck, while Oracle is seeking private credit under an active downgrade risk and Broadcom is lending its own credit rating rather than OpenAI standing on its own. These are the terms you see when lenders already doubt the borrower, so the cost of capital climbs and the structure gets reworked before it closes. The opposite outcome, all three closing cleanly at the reported size and terms, requires private-credit appetite for AI infrastructure to stay at a peak through a quarter when at least one credit signal is already flashing, which is the less likely path when the CDS market is doing the doubting for you.
Revisit by 2027-02-28: We're right if any of the three deals is publicly reported as reduced in size, pushed past its stated timeline, or closed at higher spreads than first floated. We're wrong if all three close at or near their reported size, timing, and terms.
One more thing the routing story settles. When xAI builds its own frontier model and still routes Grokbot to Claude, the lesson for buyers is that your vendor's own product team has already conceded no single model wins. Design for the swap now, while it's cheap to do.
Also covered this issue
-
Dario Amodei Calls for AI Capability Slowdown; Altman and Musk Agree
semianalysis
Three AI CEOs announced a voluntary slowdown with no enforcement mechanism, but your API costs and model capabilities won't actually change.
-
AI Leaderboard Arena Raises $200M at $3.1B Valuation
techcrunch-ai
A startup's $3.1 billion valuation now hinges on whether its crowd-voted rankings become the standard your company uses to pick which AI model to deploy and trust.
-
Fired OpenAI safety researchers deny misconduct, warn of chilling effect
techcrunch-ai
Fired safety researchers warn that OpenAI now punishes external safety review, quietly weakening the oversight you rely on without knowing it.
Comments