Refacto AI

Podcast episode

AI Model Month Is Off to a Blistering Start

agents cost-compression evals model-pricing security

Nine AI model releases dropped in a single week, and the episode hosted by Nathaniel Whittemore covers what that actually means for anyone building or buying with AI. The headline number is less interesting than what's underneath it.

The substance is benchmark rot. Google Gemini 3.8 Flash and Meta's Muspark 1.3 both post strong scores on Terminal Bench 2.1, then collapse on Terminal Bench 4.0, a version released just two weeks earlier and too new to have been trained against. Semianalysis labeled both "benchmaxed," meaning tuned to ace the public test rather than do real work. Whittemore and Tristan Buckmaster also dig into the data question: OpenAI CRO Mark Chen confirmed the company uses de-identified user feedback to improve its models, which prompted NYU's Levent Alpöge and Buckmaster to raise concerns about proprietary R&D leaking through standard API tiers.

The cost curve is real. Muspark 1.3 runs a coding task for 55 cents, about a quarter of Opus 5. But a task that fails and retries three times costs more than the sticker. Stop routing work by leaderboard position and get a written zero-retention agreement before you put anything sensitive into a standard API tier.

Full analysis

Nine model releases in a week, and the loudest signal isn't any single one. It's that the benchmark scores you use to pick a model just stopped being trustworthy. Google Gemini 3.8 Flash and Meta's Muspark 1.3 both post frontier-level coding numbers on Terminal Bench 2.1, then fall off a cliff on Terminal Bench 4.0, a version released two weeks earlier. Semianalysis called both "benchmaxed," meaning tuned to ace the public test rather than the actual work. That's the through-line: capability is getting cheaper and faster, and the way you verify capability is getting worse.

The rest of the episode is the same story from different angles. Cheaper agentic models. A personal agent from Meta. A lawsuit about whether you get the usage you paid for. And a nasty question underneath the OpenAI math claim: does the lab you build on get to see and act on your work?

What's actually being decided here. Not "which model wins." That churns weekly now. The real decision for anyone buying or building with AI is whether you keep trusting public leaderboards to route your workloads, and whether you keep sensitive work inside a vendor's consumer or standard API tier. Both are easy to undo if you move now, expensive to undo if you've wired a year of pipelines around a leaderboard number or leaked proprietary R&D into a shared training pool. Nothing here sets a hard deadline. But the benchmark rot and the data question both compound the longer you ignore them.

The Skeptic. A model that scores 89.4% on Terminal Bench 2.1 and 19.1% on Terminal Bench 4.0 is telling you exactly one thing: it was trained to pass the test everyone can see. Opus 5 dropped from 74% to 51.8% on the same jump. That's a real model degrading gracefully. Gemini 3.8 Flash collapsing to 19% is a model that memorized the answer key. And notice the pricing sleight: Artificial Analysis calls Gemini "cheapest at its intelligence level" while per-token pricing didn't change. The "40% cheaper" is fewer tokens used, not a lower rate. Read the mechanism before you believe the headline.

The Researcher. The useful distinction this week is between two benchmarks two weeks apart. Terminal Bench 4.0 is a signal only because it's too new to have been trained against. That's the whole game now. Any public eval decays into marketing the moment labs can scrape it. Meta's own Chief AI Officer Alexander Wang was refreshingly straight: Muspark 1.3 isn't as strong as the top models, it's cheaper, and future models will compete on capability. Believe that framing over the DeepSUI score. The 75.4% coding number that "edges out Opus 5" is the part that won't survive contact with your actual codebase.

The Compute Pragmatist. Ignore the leaderboard noise and the cost curve is real. Muspark 1.3 runs a coding task for 55 cents, about a quarter of Opus 5. Gemini 3.8 Flash outputs roughly 20% more tokens per second than anything close and burns fewer tokens per task. For high-volume agentic work, where the model chews through many steps autonomously, that's a genuine shift in what you route where. The move is tiering: cheap-and-fast for bulk, expensive-and-smart for the hard 10%. Just don't let the 55-cent number seduce you into sending work the cheap model can't actually finish. A task that fails and retries three times isn't 55 cents.

The Enterprise Buyer. The Navier-Stokes fight is the part that should make a CTO put down their coffee. OpenAI CRO Mark Chen admitted the company uses "de-identified user feedback to improve ChatGPT and Codex in a holistic way, and so does every LLM company." NYU's Tristan Buckmaster and Anthropic's Levent Alpöge claim they spent a year on that math problem inside OpenAI's Codex before OpenAI announced a solution. Take the innocent read and it's still a warning: your proprietary work, run in a standard tier, feeds the model that your competitor also uses. If you handle legal, financial, or R&D work, you need a written data processing agreement and a zero-retention tier. Get the contract in writing, and make sure it actually specifies retention terms.

The Builder. Meta Muse is the thing I'd actually poke at Tuesday morning. Each user gets an isolated virtual machine, and a separate "Sentinel" checks every action before it leaves the sandbox, so Muse never touches raw passwords or card numbers. That architecture is the right answer to the question every agent vendor dodges: what happens when it tries to spend your money wrong. a16z's Olivia Moore flagged the real tension. The native app connectors make it more reliable than browser-based agents, but Meta's trust deficit cuts against handing it your inbox and credit card. Bosworth says he relies on it more than anything Meta shipped recently. Founders love their own launches. Watch the retention numbers. The CTO quote is the easy part.

Where the council splits

Two real disagreements.

The Compute Pragmatist sees a cost frontier collapsing and wants to route bulk work to 55-cent models today. The Skeptic and Researcher say those same models can't be trusted precisely because their headline numbers are gamed, so you don't actually know which tasks they can finish. Both are right, and the resolution is the same: you cannot outsource that judgment to a leaderboard anymore. You have to test on your own work.

The Builder is excited about Muse's sandboxed architecture; the Enterprise Buyer is looking at the Navier-Stokes admission and getting nervous about handing any consumer-tier agent real access. The bridge is that the security question ("can the agent misspend my money") and the data question ("does the vendor learn from my work") are different problems, and Muse solves the first while doing nothing for the second.

What this actually hinges on

Three beliefs, and only you can settle the first one for your stack.

Does the cheap model finish YOUR tasks, or does it only finish the benchmark's? Build a small private eval from your real workflows and run Gemini 3.8 Flash, Muspark 1.3, and Opus 5 against it. The public gap between them is now worthless. Your private gap is the only number that routes budget.

Is your sensitive work sitting in a tier that trains the model? Check which API tier your R&D, legal, and financial prompts run through, and whether you have zero-retention in writing. Chen said the quiet part out loud. Assume every lab does the same until your contract says otherwise.

The council leans clearly: the capability gains are real and cheap, the verification is broken, and the data exposure is under-priced. Move on your own evals and your own contracts. Don't move on the leaderboard.

Prediction: Within six months, by the time Terminal Bench 4.0's coding results have their next major leaderboard refresh, at least one of Gemini 3.8 Flash or Meta Muspark 1.3 will show a gap of 25 points or more between its Terminal Bench 2.1 score and its Terminal Bench 4.0 score, confirming the benchmaxing charge holds rather than fades.

Confidence: Medium. The gap is already visible; the risk is the benchmark itself getting gamed before the comparison settles.

Why: Gemini 3.8 Flash already posts 89.4% on Terminal Bench 2.1 and 19.1% on Terminal Bench 4.0, a 70-point collapse, while Opus 5 holds a graceful 74% to 51.8%. That pattern comes from optimizing against the public test everyone can scrape, and labs shipping three Flash updates in six weeks have every incentive to keep chasing the visible scoreboard. The opposite outcome, the gap closing, would require Google or Meta to retrain specifically for the newer test and publish it, which erases the cost advantage that is their whole pitch. The one thing that breaks this call is Terminal Bench 4.0 itself getting scraped and gamed fast enough that everyone's scores rise together, which would muddy the comparison rather than clear it.

Revisit by 2027-03-15: We're right if a public Artificial Analysis or Semianalysis comparison shows either model with a 25-plus point spread between Terminal Bench 2.1 and 4.0. We're wrong if both models close the spread to under 15 points on the newer test.

Comments