Industry story
Anthropic Releases Claude Haiku 5.5 at Sharply Lower Price
agents cost-compression inference model-pricing
Anthropic's 10x price cut on Claude Haiku 5.5 ($0.10/$0.50 per million tokens, down from $1/$5) lands softer than it looks. A tokenizer change inflates your word count by roughly 25%, mandatory chain-of-thought reasoning burns compute on every call whether you want it or not, and the price parity with OpenAI's GPT-6 Luna evaporates the moment your prompts cross 100,000 tokens. The real cut is closer to 7-8x for short workloads, and a wash or worse for long-context and agent jobs. Measure your actual corpus and your P95 latency before you reprice anything.
Full analysis
Anthropic cut the price of its cheapest model, Claude Haiku, by 10x overnight: $0.10 per million words in, $0.50 per million out, down from $1/$5. That is the headline. The story is what Anthropic buried under it. The new model forces on an internal "thinking" step (the model talks itself through the problem before answering) that you can't turn off, uses a word-counting scheme that charges you about 25% more words for the same text, and gets 5x more expensive the moment your prompt crosses 100,000 tokens. The sticker says cheap. The receipt says something narrower.
This is easy to undo. Swapping a cheap model in and out of a routing or classification job is a config change, not a rewrite. What's hard to undo is signing a pricing contract or committing a product margin based on the sticker price before you've measured the real one. No deadline forces your hand. The old Haiku 4.5 still runs. Nobody is making you move this month.
The Skeptic. A 10x cut sounds like Anthropic blowing up the floor. It's Anthropic catching up to where OpenAI already parked GPT-6 Luna. Same $0.10/$0.50, and the quote says so plainly. The tokenizer change is the quiet part: a worse word-counter that adds ~25% to your bill on the same prompts, so the real cut is closer to 7-8x. Then the mandatory thinking step burns extra compute on every single call, which is how Anthropic claws the margin back. And the parity with Luna evaporates at exactly 100,000 tokens, where the long-document and agent workloads actually live. Cheap where you don't need it, expensive where you do.
The Compute Pragmatist. Per-token pricing is the magician's hand here. Every call now runs a thinking trace at "medium effort" by default, so each request chews materially more compute than the old straight-through model, even at a lower price per word. The 5-minute ceiling on max-effort runs tells you this isn't open-ended search. It's a fixed-depth scratchpad with a hard token budget. The labs keep teasing think-as-long-as-it-needs scaling; this is not that. The new word-counter also reshuffles memory and throughput on Anthropic's own servers. Net: the cost-per-finished-task barely moved from where it looks on the price page, because you're paying for more words and more thinking to get the same answer.
The Builder. Tuesday morning, this ships for high-volume routing, classification, and light extraction that were marginal on the old price. That part is real and immediate. Two things break first. The always-on thinking step means anything expecting sub-second responses needs a load test before cutover, because the demo latency won't match your tail latency under mandatory reasoning. And re-run your token counts on your actual production text before you quote anyone a price, or the 25% word-counter inflation becomes a budget surprise after you've committed. Keep your calls under 100,000 tokens by design. Cross that line and Luna is the better buy, full stop.
The Safety Lens. Forcing the thinking step to always be on, with no off switch, is the most consequential choice in the release. In principle it helps: the model writes out its reasoning before it acts, which makes its behavior easier to audit and harder to hijack mid-task. In practice, the open question is whether that reasoning trace is walled off from the tools the model can call, or whether a clever adversarial prompt can steer the think-out-loud step into talking itself into a bad action. Anthropic has not published the thinking-mode safety evals. At 3.4 cents per max-effort run, this lands inside millions of cheap agent loops, so the total risk surface grows even as each call gets cheaper.
Where they disagree. The Builder sees a deployable win at the sub-100k price; the Skeptic sees parity with a competitor plus two hidden surcharges. Both are right, and the gap between them is the whole decision: the headline cut is real, the delivered cut is narrower, and how much narrower depends entirely on your own text and your own latency budget. The second split is the Compute Pragmatist versus the price page. Anthropic lowered the number you see and raised the compute you spend, which means the per-task cost moved far less than 10x for most real jobs. The Safety Lens adds the part nobody can resolve yet: mandatory visible reasoning looks like interpretability, but looks-like isn't proven-under-pressure, and the evals aren't out.
What this hinges on. Three facts, all checkable on your own machine. Does the effective price, after the 25% word-counter inflation and the forced thinking tokens, actually beat your current setup on your corpus? Does the always-on reasoning blow your latency budget at P95? And does your workload stay under 100,000 tokens, or does the 5x cliff eat the savings? The council leans skeptical on the headline and practical on the use: a genuine win for short, high-volume, latency-tolerant jobs, a wash or worse for long-context and agent work where Luna's cheaper tiers win. Before you move anything, count your tokens on real production text and run a latency test with reasoning on. The sticker is not the bill.
Prediction: Within 60 days of Claude Haiku 5.5's launch, Anthropic will ship a way to reduce or disable the forced reasoning step on Haiku 5.5 (a minimum-effort or off setting), responding to latency-sensitive developers, by 2026-12-07.
Confidence: Medium. The always-on constraint directly breaks the cheap, high-volume jobs Haiku exists to serve.
Why: Haiku's entire reason to exist is cheap, fast, high-volume work: routing, classification, extraction, where sub-second response time is the whole point. A mandatory thinking step that can't be turned off adds latency and compute to every one of those calls, which is exactly the workload that cares most about both. Anthropic already lets you dial reasoning effort on its larger models, so the control exists in their stack and withholding it on the cheap model fights the product's purpose. The opposite outcome, Anthropic holding the line on forced reasoning, only makes sense if the model is unreliable without it, and a lab that just cut price 10x to win volume will not want to bleed latency-sensitive customers to a competitor whose cheap model answers faster.
Revisit by 2026-12-07: We're right if Anthropic documents a way to set Haiku 5.5's reasoning to minimum or off via the API. We're wrong if the medium-effort default remains the floor with no lower setting.
Comments