Refacto AI

Podcast episode

What to Use the Latest AI Tools For

coding-agents cost-compression inference model-pricing open-weights

Nathaniel Whittemore's latest episode works through what the current crop of AI tools is actually good for, using a week of notable releases as the test cases. OpenAI paused new $200-a-month Pro subscriptions because it couldn't meet demand. At the same time, DeepSeek's V4.1 Flash model scores level with much pricier alternatives at $0.30 per million input tokens, and Cognition's SWE-2 hits 50% on a hard coding benchmark at 64% less cost than its predecessor. The cheap tier stopped being a compromise this week.

Jensen Huang's hardware numbers explain why the top end is tightening: rack-scale GPU systems now run $8.5 million a unit, and Microsoft is tripling its data-center power capacity by 2032. When the infrastructure is supply-constrained and buyers are spending at that scale, flat-rate frontier access doesn't survive.

The practical move is simple. Swap your batch coding and data-pipeline work to cheaper models now. You can change a model endpoint in an afternoon, and the direction on pricing is already set.

Full analysis

Two things in this episode point the same way, and both cost you money. OpenAI paused new $200-a-month Pro subscriptions because it can't feed the demand, and analysts think that tier never comes back. Meanwhile the cheap tier just got good: DeepSeek's V4.1 Flash beats DeepSeek's own flagship at a quarter the price, and Cognition's SWE-2 hits 50% on a hard coding test at 64% less cost than the model it replaces. The subsidized-token party is ending at the top and the cheap seats are getting comfortable.

What's actually being decided: whether you keep routing production work to premium frontier models on a flat subscription, or move the bulk of it to cheap-and-good models before the pricing gets restructured out from under you. This is easy to undo. You can switch model endpoints in an afternoon. The deadline isn't fixed, but the direction is: OpenAI already stopped taking new Pro signups, so the clock is running whether or not there's a date on it.


The Skeptic

Watch the benchmark sleight before you re-architect anything. "SWE-2 scores 50% on FrontierCode 1.1, ahead of GPT-5.6 Sol" tells you nothing about your codebase. These coding tests are narrow, and every lab tunes to them. Cognition's own testers reporting "6× more merged PRs" is a vendor number from a private preview, which is the least trustworthy kind. And "the era of subsidized tokens is ending" is Matt Schumer, an analyst, selling a vibe. OpenAI paused ONE tier. That's a capacity story, not a confirmed price hike. Don't rebuild your stack around a paraphrased warning and a leaderboard you can't reproduce.

The Compute Pragmatist

The hardware numbers are where this episode stops being speculative. Jensen Huang says a rack-scale GPU system is now $8.5 million a unit, 2 million parts, and orders are growing 27% month over month. Microsoft is going from 12 gigawatts of data-center power to 38 by 2032, with the AI slice rising from 2 to 13. That is not the shape of an industry about to cut prices. When the people selling the picks and shovels are supply-constrained and the buyers are tripling capacity, the token you rent gets more expensive. OpenAI pausing Pro signups is the first crack. Plan your unit economics as if the flat-rate buffet is closing.

The Open-Source Advocate

This is the week the cheap tier stopped being a compromise. DeepSeek V4.1 Flash: 552 billion parameters, $0.30 per million input tokens, $1.20 output, scoring level with GPT-5.6 Luna and Gemini 3.8 Flash on a public intelligence ranking. Cognition's SWE-2 is post-trained on Kimi K3, a Chinese open model, and it's competitive on coding. The pattern is clear: the good-enough models increasingly come from open Chinese weights you can run or rent cheaply, while the frontier US labs ration access. If your workload is batch coding or data-pipeline work, not real-time customer chat, you are paying a premium for a logo.

The Safety Lens

Anthropic's misuse report is the uncomfortable subplot. Moonshot AI routed roughly 300,000 customer requests through Claude's Opus model over ten days, apparently to harvest realistic queries for its own training. Alibaba, DeepSeek, and Xiaomi ran distillation attacks, using fake accounts to pull out the model's reasoning steps and copy them. Anthropic also blocked a grant tied to gain-of-function work on a virus for a military institute. If you build on frontier APIs, this is the argument the labs will use to justify tighter access, more identity checks, and higher prices. The scarcity and the safety story reinforce each other, and both land on your bill.

The Builder

GPT Live 1 is the one thing here you can ship on Tuesday. Full-duplex voice, meaning the model talks and listens at the same time, handles interruptions, manages background noise, and can call tools mid-conversation, at five cents a minute plus the usual backend cost. That opens real voice agents that don't feel like walkie-talkies. But everything else in this episode says: don't bet your production spend on one vendor's premium tier. Build your coding and data jobs so you can swap the model behind them. DeepSeek Flash and SWE-2 are cheap enough that a model-router setup pays for itself the first month the frontier price moves.


Where the council splits

The Skeptic and the Open-Source Advocate genuinely disagree on the coding benchmarks. One says the leaderboard is gameable and vendor-reported; the other says the price gap is so large that even a discounted version of the claim wins. Run your ten hardest real production tasks through both the cheap and the premium model and grade the output yourself. That settles it faster than any leaderboard.

The bigger tension is between the Compute Pragmatist and anyone hoping prices fall. The hardware math says frontier tokens get scarcer and dearer. The open-model math says the floor is collapsing. Both are true at once: the top gets pricier and rarer, the middle gets cheaper and better, and the gap between them widens. That's the actual decision. Not "which model is best," but "which of my workloads actually needs the frontier, and which am I overpaying for out of habit."

What this hinges on

One belief: that cheap open models are now good enough for your non-frontier work. You can settle it this week. Take your ten hardest real production tasks, run them through DeepSeek V4.1 Flash and SWE-2 alongside your current premium model, and grade the output yourself. If the cheap tier holds up on eight of ten, move that traffic and keep the frontier for the two that need it. Build the router now, while switching is an afternoon and not a fire drill.


Prediction: Before OpenAI's next major consumer pricing update (expected alongside or shortly after the GPT-6 Astra general rollout), OpenAI will not reinstate a flat $200/month Pro tier with unlimited or unthrottled frontier access; access at that price point will be replaced by usage-metered or capacity-gated tiers.

Confidence: Medium. Supply is the binding constraint and metered billing follows scarcity.

Why: OpenAI paused new $200 Pro signups specifically to protect existing users, which is what a company does when demand outruns compute and a flat all-you-can-eat price is losing money on the heaviest users. Huang's $8.5M rack units and 27%-month-over-month order growth, plus Microsoft tripling data-center power, all say frontier compute stays scarce and expensive for years, so the cost per heavy Pro user doesn't fall fast enough to justify bringing the flat tier back. The opposite outcome, OpenAI restoring unlimited flat pricing, would require either a sudden supply glut or a decision to keep subsidizing power users indefinitely, and nothing in this episode points that way. The one thing that saves the flat tier is a compute breakthrough that undercuts the scarcity story, which is the less likely bet given the hardware order book.

Revisit by 2027-03-17: We're right if OpenAI's Pro-level access is sold on metered, tiered, or capacity-limited terms rather than a flat $200 unlimited plan. We're wrong if a flat $200/month Pro tier with unthrottled frontier access is generally available to new customers.

Comments