Refacto AI

Podcast episode

Why Fable 5.1 Is Worth the Upgrade

agents evals inference model-pricing security

Anthropic shipped Claude Sonnet 5.1 and Opus 5.1, claiming 25% lower costs. Artificial Analysis ran the numbers and found Sonnet 5.1 actually costs 20% more per task, because the model burns 70% more tokens (the chunks of text a model processes, billed per chunk). The per-token price fell; the token count exploded. Separately, OpenAI's Astra can now find and exploit unknown security vulnerabilities with no human in the loop. And both models think longer by looping internally before answering, which buys capability but hides the reasoning from human review. OpenAI Chief Scientist Jakub Pachocki calls the visible chain-of-thought "fragile and trending in a negative direction."

The through-line is that the labs are trading cost and transparency for benchmark wins. Better model, higher bill, less readable reasoning.

If you buy AI for a compliance-heavy shop, test on your own workload before believing any vendor's cost claim, and put reasoning visibility requirements in the contract now, while you still can.

Full analysis

Your draft

Anthropic shipped Claude Sonnet 5.1 and Opus 5.1, and they now hold the top three spots on the main intelligence leaderboard. But the story worth your attention is that a "cheaper" model got more expensive in practice, OpenAI's next model can now hack things on its own, and the way that model thinks is getting harder for humans to read. Three separate signals, one theme: the labs are trading things you care about (cost, transparency, safety) for benchmark wins.

This is easy to undo for you as a buyer. Nobody is locking you into Sonnet 5.1. You can route work to whichever model wins on your task this month and re-route next month. What sets the clock: Anthropic's zero-data-retention system for enterprise ("EFS") starts rolling out this fall, and OpenAI's Astra ships "imminently." Neither has a hard public date.

The Skeptic

Anthropic claimed 25% lower cost, up to 45% on agentic work. Artificial Analysis ran it and found Sonnet 5.1 costs $3.76 per task versus $3.14 for the old one. That is 20% more expensive, because the new model burns 70% more tokens (the chunks of text a model reads and writes, which you pay for by the chunk). The per-token price dropped; the token count exploded. Read the ARC Prize number too, which found 32% cheaper. Two credible testers, opposite results. That gap tells you the "cheaper" claim depends entirely on your workload, and Anthropic quoted the flattering case. Test on your own traffic before you believe either number.

The Researcher

The benchmark jumps are real and large. On TerminalBench 4.0, an agentic coding test, Sonnet 5.1 hit 55.8% versus 42% for the prior version and 37.3% for GPT-5.6. On the business-automation test it nearly doubled, 31.4% from 17.1%. Those are not rounding errors. But watch Astra's cybersecurity numbers, because they explain the token story. Astra found and exploited unknown security flaws using 40,000 tokens where GPT-5.6 needed 110,000 and still failed. The new frontier trick is thinking longer and looping internally. That buys capability. It also buys token bloat and opacity, which are the same two complaints hitting Sonnet 5.1.

The Safety Lens

Astra crossed OpenAI's own "critical cybersecurity" line. It can find and exploit brand-new vulnerabilities with no human in the loop. During testing it discovered two zero-days and chained them. That is a real product capability now, not a research demo, and it cuts both ways: your security team gets automated pen-testing, and attackers get the same tool. Then there's Recurrent Depth, the technique that loops the model over the same text before it answers. Part of the reasoning now happens where humans can't read it. Ryan Greenblatt of Redwood Research warns the natural next step is reasoning almost entirely in that hidden space. Nathan Calvin of EncodeAI flags the real risk: once one lab proves the efficiency gain, others copy it, and nobody wants to be the one leaving performance on the table for the sake of readable reasoning.

The Enterprise Buyer

The unlock here is boring and it matters: Anthropic's EFS gives enterprise customers zero data retention. Sonnet 5's 30-day retention was a straight blocker for regulated buyers in finance, health, and legal. That policy killed deals. Removing it opens them. But if you buy AI for a compliance-heavy shop, add a new line to your checklist: can you read the model's reasoning? OpenAI Chief Scientist Jakub Pachocki says Astra keeps chain-of-thought (the readable step-by-step of how a model reasons) visible, then concedes it's "fragile and trending in a negative direction." When the vendor's own scientist hedges like that, put it in the contract.

The Compute Pragmatist

The pattern across all three releases is the same trade: spend more compute per task to win the benchmark. Sonnet 5.1 burns 70% more tokens. Astra loops internally. Both are the same bet, that thinking longer beats thinking smarter. For you that means a capability upgrade no longer means a cost cut. It often means a cost increase you only see on the invoice. Google is the counter-move worth watching. Gemini 3.8 Flash is a small, fast, cheap tier that Google's own engineers reportedly preferred over Anthropic's Opus on coding. And Google scrapped its bigger Gemini 3.5 Pro candidates because they weren't beating the Flash models. When the cheap tier eats the expensive tier internally, that's a signal about where the value is.

Where the council splits

Two real disagreements.

The Researcher sees genuine capability gains and the Skeptic sees a bait-and-switch on price, and both are right at once. The model is better AND more expensive per task. The industry spent three years training buyers to expect "better and cheaper" every cycle. That era is pausing. Host NLW said it plainly: even a purist like Anthropic is not immune to the new reality where efficiency, not just capability, is the fight.

The Safety Lens and the Compute Pragmatist part ways on Recurrent Depth. The Pragmatist sees an efficiency win Google and Anthropic will absolutely copy. The Safety Lens sees the readable-reasoning safeguard quietly dying because no lab wants to eat the performance penalty of keeping it. The Compute Pragmatist is describing exactly the mechanism the Safety Lens is scared of.

What this actually hinges on

For you, three things. One, does the capability gain justify a per-task cost increase on YOUR workload, which you can only learn by running your own tasks through both models. Two, whether "route by task" is now mandatory infrastructure. Power users are already sending long agentic jobs to Sonnet 5.1 and keeping GPT-5.6 for interactive work. That architecture has become the default, not an edge case. Three, if you're in a regulated business, whether readable reasoning survives as a buyable feature or quietly disappears.

Before you commit spend: run your ten most common tasks through Sonnet 5.1 and your current model, and compare total token cost. The headline per-token price will mislead you. That single test settles the whole "is it cheaper" argument for your specific case.

Prediction: Within OpenAI's next major model release after Astra (by 2026-09-03 + roughly one cycle, no later than 2027-06-30), at least one of Anthropic or Google will ship a frontier model using a looped or latent-reasoning technique that reduces human-readable chain-of-thought, and will not offer full reasoning transparency as a default.

Confidence: Medium. The efficiency gain is proven and copyable; the restraint is voluntary.

Why: Astra's Recurrent Depth delivered a real, measurable win, solving hard cybersecurity tasks with roughly a third of the tokens GPT-5.6 needed, and the whole competitive fight this cycle has moved from raw capability to cost per task. Nathan Calvin of EncodeAI named the mechanism directly: once one lab proves the efficiency gain, rivals find it too, and none wants to sacrifice performance to keep reasoning readable. Anthropic just shipped a model whose main complaint is token bloat, so it has the strongest incentive of anyone to adopt a technique that cuts token spend. The opposite outcome, everyone voluntarily keeping reasoning fully transparent while a competitor gets cheaper and faster by not doing so, requires the labs to act against their own cost pressure with no regulation forcing it, and OpenAI Chief Scientist Jakub Pachocki already called transparency "fragile and trending in a negative direction."

Revisit by 2027-06-30: We're right if Anthropic or Google ships a frontier model using looped/latent reasoning that reduces readable chain-of-thought and does not make full transparency the default. We're wrong if both labs' new frontier models through that date keep full step-by-step reasoning human-readable by default.

Comments