Refacto AI

Podcast episode

The Most Important Trends Showing Up in New AI Products

agents cost-compression inference model-pricing orchestration

"The Most Important Trends Showing Up in New AI Products" is a roundup from Nathaniel Whittemore and Dillon Rolnick covering the week's AI product moves. The headline: cheap frontier-quality models are now American, and even the big closed labs are routing users to whichever model wins the task rather than defending their own.

The specific numbers matter here. Anthropic's Haiku 5.5 (a small, fast, cheap model) hits 72.4% on OSWorld, a benchmark for whether a model can drive a computer by clicking and typing, against 48.9% for GPT-6 Luna. Real-task cost runs 12 to 21 cents. Rolnick's Nous Research, an open-source model shop, is tracking to $100M revenue and just hit a $1.5B valuation. Meanwhile, SpaceX is shopping a two-page deck to raise $40B for NVIDIA chips, and its credit default swaps price a 15% default chance by 2031.

The cheap-inference story is real, but the cost floor is being held up by leveraged financing. Those two things are in conflict.

Analysis

Showing the shorter version.

Raw Model Capability Is No Longer the Competition

Two things happened this week that matter more than the billion-dollar headlines. The cheapest frontier-quality models are now American. And even the labs with the best models are routing you to whoever wins the task. Both point the same direction: the model underneath is a commodity. The price and the product are what matter.

The Anthropic Haiku 5.5 numbers are specific. 72.4% on OSWorld, a benchmark for whether a model can drive a computer by clicking and typing, against 48.9% for GPT-6 Luna. More than double GPT-6 Luna on TerminalBench 4.0. Those are real gaps on tasks operators actually run. The pricing is the louder signal: 25% of Sonnet's per-token cost, with an 80% discount on prompts under 100K tokens. Artificial Analysis pegs real task cost at 12 to 21 cents per task. Haseeb Qureshi's framing, "buy American for the cheapest LLMs," is a genuine reversal. DeepSeek's price edge is gone.

The routing story confirms what the pricing implies. xAI built a frontier model and still routes Grokbot to Claude and Midjourney. Elon Musk's own product team is conceding no single model wins everything. The RAMP data, which tracks what companies actually pay for AI each month, shows OpenRouter and Featherless, both model routers, showing up in enterprise spend. Operators are already voting for optionality with budget. The right builder response is to stop writing code against one vendor's SDK and put a router in front, so swapping Haiku for whatever beats it next month is a config change.

Nous Research hitting a $1.5B valuation reinforces this. 22 million installs, roughly 2.5% of all tokens used globally, tracking to $100M revenue by year-end. CEO Dillon Rolnick's pitch is observability and no lock-in. That pitch is landing precisely because the routing trend proves his point.

OpenAI's Intelligent UI, charts and interactive output rendered inline, gets attention. It's trivially copyable. Anthropic or Google ships the same thing in a quarter. If you've built your own front end, it does nothing for you. If you ship inside ChatGPT, it just raised the bar on what users expect. Not nothing, but not a moat.

The part the cheap-inference story hides is the financing underneath it. SpaceX is reportedly shopping a two-page deck to raise $40B for NVIDIA chips, and its credit default swaps are pricing a 15% chance of default by 2031. Oracle is seeking private credit under downgrade risk. Broadcom is lending its own credit rating to OpenAI's roughly $50B custom-chip build. Haiku at 12 cents a task looks like deflation, but the hardware underneath is financed like a leveraged buyout. If any of these deals wobble, spot inference prices do not stay this low.

That is the real disagreement. The cheap-inference camp sees a genuine price collapse: American labs closed the gap, open models have a business, routers let you always buy the cheapest winner. The skeptical camp sees low prices propped up by junk-grade debt and says the floor is borrowed. Both cannot be right for long.

The call: At least one of the three major AI chip financing deals (SpaceX's $40B NVIDIA raise, Oracle's private-credit chip purchase, or Broadcom's credit-backed OpenAI silicon deal) gets downsized, delayed past its reported timeline, or repriced at materially worse terms before Q1 2027 earnings land in February 2027. Medium confidence. The debt is already being priced at distress levels. The opposite outcome requires private-credit appetite for AI infrastructure to hold at a peak through a quarter when the CDS market is already doing the doubting.

The practical question is whether 12-to-21-cent frontier inference is a stable fact to build on or a promotional price held up by debt. Run the test. Rebuild one real workload on Haiku 5.5, measure the task cost yourself, and put a router in front so you are not married to it. If the price is real, you win. If it's borrowed, you're hedged.

Also covered this issue

Comments