Podcast episode
Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models
agents evals inference model-pricing tool-use
Nathaniel Whittemore's show this week covered four AI model releases and two regulatory moves that all landed in the same few days: Anthropic's Claude Sonnet 5.5, Google's Gemini 4 Argon (announced but not yet released), an FTC investigation into OpenAI and Anthropic over rogue agents (AI that takes actions on your behalf, like placing orders or filling forms), and a White House voluntary safety accord signed by the major labs.
The substance that matters is the gap between benchmark scores and what you actually pay per task. Artificial Analysis found Sonnet 5.5 costs 27% more per task than Opus 5.5 at max inference settings, despite Anthropic's advertised 30% savings. Token appetite is the variable the headline price hides. Builder Kun Chen's approach, routing Opus to plan and Sonnet to implement, points where this is going: you pick a model per step, not a winner.
Gemini 4's benchmark numbers look strong. Google's own employees told Bloomberg it chokes on real coding. You can't test it yet anyway. Wait.
Full analysis
Four models, two regulators, and one White House photo op landed in the same week. For anyone running AI in a business, the signal worth pulling out is simple: the gap between what a model scores on a benchmark and what it costs you in production is widening. The price you pay per task is now the thing that moves.
What's actually being decided: nothing you have to decide today. Gemini 4 Argon isn't released. Sonnet 5.5 is, and it's the one live choice in front of you. The White House accord and the FTC probe set the weather for anyone deploying agents (AI that takes actions on your behalf, like placing an order or filling a form), but neither forces a move this month. How hard is this to undo? Swapping a model behind an API call is easy to undo. Rewiring a multi-agent pipeline to tune token spend is harder. What sets the deadline? Gemini 4's 50% launch discount is the only clock, and you can't even use the model yet, so there's no real deadline here.
The Skeptic. Two benchmark claims blew up on contact this week, and they're the whole story. Anthropic says Sonnet 5.5 is 30% cheaper. Artificial Analysis ran it and got $7.60 per task at max settings, 27% more expensive than Opus 5.5 and more than double GPT-6 Astra. "Cheaper per token" and "cheaper per task" are not the same sentence when the model burns more tokens to finish the job. Google's Gemini 4 tops the Vals index, while its own employees tell Bloomberg it chokes on real coding. We have seen this movie with Gemini 3 Pro. A benchmark you can't reproduce in your own workload is marketing.
The Researcher. Look at where each model wins and you learn what it's for. Sonnet 5.5 jumped from 10.3% to 70.6% on Terminal Bench 4.0, a test of running commands in a terminal, and beat Opus on it. That is a real, specific gain in tool use, not a vibe. Gemini 4 tripled the field on Harvey's legal agent benchmark but sits 10 points behind the leader on Frontier SWE coding. No single number crowns a winner anymore. The Intelligence Index, a blended score Artificial Analysis publishes, puts Opus 5.5 first and Sonnet 5.5 second at 56. Useful as a starting filter. Useless as a deployment decision.
The Builder. The deployment pattern that matters showed up quietly. Builder Kun Chen's stack: Opus to plan, Sonnet to implement, Fable as the fallback. Three models from one lab, each doing the part it's cheapest and best at. That's where this is going. You don't pick "the best model," you route each step to the model that does it well. Which means your real work is building a test harness that measures cost and quality per step, then tuning inference settings. Artificial Analysis found cutting Sonnet 5.5 from "max" to "extra high" dropped cost by two-thirds. If you run it at max out of the box, you're lighting money on fire.
The Safety Lens. The FTC opened a real investigation into OpenAI and Anthropic over rogue agents, triggered by the Hugging Face incident and dozens of others, and it's drafting civil investigation demands, which are subpoenas by another name. The legal theory is bipartisan: existing product safety law already covers what your agent does, no new AI statute required. Lina Khan and David Sacks agree on that, which almost never happens. The White House accord is "morally binding," which is Trump's own word for voluntary. If you deploy consumer-facing agents, the compliance surface is the actions your agent takes, and nobody is waiting for a new law to come after it.
The Compute Pragmatist. Per-task cost is now the live number, and it's volatile. Gemini 4 at $1.99 with a 50% discount is above Astra once the discount ends. Sonnet 5.5's advertised savings vanish at max inference. The lesson: headline price per token tells you almost nothing about your bill. Token appetite does. A model that thinks harder to get the right answer can cost more than a pricier model that answers in fewer tokens. Until you've run your own workload through it, every pricing claim is someone else's benchmark, not your invoice.
Where they disagree. The Researcher sees Sonnet 5.5's Terminal Bench leap as a genuine capability win. The Compute Pragmatist and Skeptic say that win is worthless if it costs you 27% more than Opus to collect it. Both are right, and that tension is the decision: Sonnet 5.5 is a strong sub-agent for implementation steps inside a pipeline you've tuned, and a bad standalone default if you run it at max and never check the token count. The second split is Safety versus the room: Douglas Farrer gave the FTC probe "0.0%" credibility the same week the lab CEOs posed with Trump. The probe may be theater. The legal theory underneath it is not, and that's the part that outlives the photo.
What this hinges on. One belief: does Gemini 4 hold up in production the way it holds up on the Vals index? Google's own employees say no, Google says yes, and you can't test it because it isn't released. So don't redeploy anything to it. On Sonnet 5.5, the question is answerable right now. Run your actual workload through it at "extra high," not max, measure cost and quality per step against Opus 5.5, and only substitute where it wins on both. On agents, start writing down your safety controls, because the demands going to OpenAI and Anthropic will set disclosure expectations that roll downhill to everyone deploying their models.
Prediction: When Gemini 4 Argon opens to Google Ultra subscribers and independent testers run it, at least one major third-party evaluator (Artificial Analysis, Vals, or SWE-bench) will report it underperforming its announced coding benchmarks on real production coding tasks, by the first full independent eval published after general API access opens.
Confidence: Medium. Google's own employees already flagged it; the Gemini 3 Pro pattern repeats.
Why: Bloomberg reported Google's own engineers with model access say Gemini 4 "struggles to handle certain coding tasks," and Gemini 4 already sits 10 points behind Astra 6 on Frontier SWE in Google's own announced numbers, so the coding weakness is visible before anyone outside has tested it. Google has a track record of announcing strong benchmarks that soften under independent testing, which is exactly why it's holding the model back and citing cybersecurity rather than shipping it. The opposite outcome, Gemini 4 matching its benchmarks in independent production coding tests, would require the internal skeptics to be wrong and Google to break its own recent pattern in the same release it chose not to ship openly.
Revisit by 2027-04-06: We're right if a named third-party evaluator reports Gemini 4 Argon trailing its announced coding scores on real coding tasks after general API access opens. We're wrong if independent production coding evals confirm Gemini 4 at or above its announced coding benchmarks.
Comments