Refacto AI

Podcast episode

Grok 4.6 Shows How Fast Your AI Options Are Expanding

cost-compression evals gpu-supply inference model-pricing

Nathaniel Whittemore's AI Briefing this week covers the Grok 4.6 release, a new model from Elon Musk's xAI that scored 61 on the Artificial Analysis Intelligence Index, tied with OpenAI's GPT-5.6 Sol and just behind Claude Opus 5, priced at $2 per million input tokens, 60% cheaper than Sol.

The cost math is real but complicated. Per task (what you actually pay when a model reasons through a problem), Grok 4.6 runs about 32% cheaper than Sol, not 60%, because verbose reasoning eats the token discount. Musk says Grok 4.7 ships in three to four weeks and will beat everything, citing SpaceX training data. Sergey Brin is apparently back at Google. The DeepSeek V4 Pro benchmark situation is a cautionary tale: leaked numbers looked great, independent scoring put it near the bottom of its own model family.

The pricing floor is moving, but a five-point index jump probably won't show up in your actual user traffic. Run the eval; don't re-plumb your stack on a press release.

Full analysis

Grok 4.6 landed this week with a benchmark score of 61 on the Artificial Analysis Intelligence Index, tied with GPT-5.6 Sol and a hair behind Claude Opus 5, and priced at $2 in, $6 out per million tokens. That's 60% cheaper per token than Sol. For anyone running high-volume inference, the top of the model market just turned into a real three-way price fight. The question worth answering: is this a durable shift in your cost calculus, or the usual pre-release hype cycle that resets in a month?

This is a Type 2 decision for most teams. Trying Grok 4.6 on one workload is a weekend eval and a rollback, not a marriage. The forcing function is real though: OpenAI just cut prices, xAI undercut them, and Musk says 4.7 ships in 3 to 4 weeks. The pricing floor is moving under your feet.

The Skeptic. Musk says 4.7 will beat everything because SpaceX training data is "so awesome and unique." Rocket telemetry improving your customer-support agent? Show me the eval. Every lab has a better model "more or less ready to go," per NLW, held back by government pressure or caution. Maybe. Or maybe the marginal gains got expensive and everyone's managing expectations. The Grok 4.5-to-4.6 jump was 56 to 61 on one composite index. That's five points on a benchmark that compresses a dozen tasks into a single number. Before you re-plumb your stack, ask whether five index points shows up anywhere in your actual traffic. For a PM: the leaderboard moved a little; your users may not notice at all.

The Researcher. The per-task figure is more honest than per-token here. Grok 4.6 runs $0.84 a task in Artificial Analysis's harness against Sol's ~$1.24, so 32% cheaper per task even though it's 60% cheaper per token. Models that reason more verbosely eat their own token discount. The DeepSeek V4 Pro leak is the cautionary tale: 87.9% on TerminalBench 2.1 in the leak, then 53 on the independent index, one point above the Flash version. Vendor benchmarks and third-party runs diverge, sometimes wildly. Trust the harness you don't control. For a PM: the company's own numbers looked great; the neutral referee scored it near the bottom of its own family.

The Open-Source Advocate. The open-weight story is where the ceiling shows. DeepSeek V4 Pro at $1.32/$3.96 per million looked like it might pressure the frontier and didn't, scoring 53. Meta's Llama release ("Muse Glimmer") got a Treasury Secretary retweet, which is a policy signal, not a capability one. The genuinely useful news is the White House reversing course to fold open-weight models into its voluntary testing framework once they hit parity. If that framework becomes an enterprise trust badge, open models that skip it lose deals they'd otherwise win. For a PM: the free-to-self-host options still aren't matching the paid frontier on hard tasks, but the government just gave them a path to legitimacy.

The Compute Pragmatist. The cheap tokens sit on top of a supply crunch that isn't easing. CoreWeave's backlog is $104B and grew $25B after the quarter closed. Nebius's Arkady Volozh says he could sell all of 2027's capacity today. Blackwell GPUs cleared auction at 15% above the prior Hopper record. So model prices are falling while the hardware underneath gets scarcer and pricier. That gap is being funded by capital, not efficiency, and capital reprices. Cognition raised at $40B, up 50% in three months. Lovable at $13.3B. When GPU access is the binding constraint, today's $2 input price is a customer-acquisition number, not a cost-plus one. For a PM: the low price is partly a land grab paid for by investors, and land grabs end.

The Builder. Grok 4.6 on Tuesday means a config change and an eval run, and xAI's API is close enough to the OpenAI shape that swapping is cheap. Gavin Baker calls the price-performance "absolute Pareto dominance for xAI and Cursor." Fine, for the workloads where it holds. The trap is the 3-to-4-week cadence. Musk ships 4.7, OpenAI answers, Anthropic drops Opus 5.5, and if you hard-wire to one model you're re-evaluating monthly. Build a router that treats models as interchangeable behind your own eval gate. The real Anthropic story here is operational: the 30-day prompt retention for government safety checks is blocking enterprise adoption outright, which is why Ramp shows Opus 5 at just 6% of Anthropic tokens. That's a compliance blocker, not a capability one, and no price cut fixes it.

Where the council splits

Three real disagreements.

The Skeptic and the Builder part ways on whether to move at all. The Skeptic says five index points won't show up in your traffic, so don't bother. The Builder says the switching cost is a weekend, so bother anyway and let your own eval decide. Both are right, and the tiebreaker is your volume. At a million queries a day, 32% per-task savings pays for the eval in a week. At ten thousand, it doesn't.

The Compute Pragmatist and everyone celebrating the price cut disagree on whether $2 input is real. Cheap tokens on top of GPUs clearing 15% above the last record, funded by CoreWeave's $104B backlog and $40B venture rounds, is a subsidized price. The Pragmatist says enjoy it and don't build a business model that assumes it lasts.

The Researcher and the Open-Source Advocate split on the benchmark that decides the open-weight case. DeepSeek's leak said parity; the independent run said mid-pack. If you're weighing a self-hosted open model to dodge Anthropic's retention problem, whose number do you plan against? The Researcher says the one from the harness you don't control, every time.

What it actually hinges on

Two beliefs. First: does Grok 4.6's per-task advantage hold on your task distribution, not Artificial Analysis's? The per-token discount is real; whether it survives a chatty reasoning model on your long-tail queries is an empirical question only your eval answers. Second: is $2/$6 a durable price or a subsidized one? The compute supply data says subsidized. Plan renewals accordingly.

The council leans act-but-don't-commit. Run Grok 4.6 through your own harness this week on real traffic, measure cost per completed task and not per token, and keep your model layer swappable because the release cadence guarantees this repeats in a month. Don't sign anything long that assumes today's price. And if Anthropic's retention policy is what's keeping you off Opus 5, understand that's the blocker, not the model.

Prediction: By the time Grok 4.7 posts an independent Artificial Analysis Intelligence Index score (Musk's stated 3-to-4-week window, so by mid-September 2026), it will not top both GPT-5.6 Sol and Claude Opus 5 on that index, despite Musk's claim it will "exceed all current models."

Confidence: Medium. The 4.5-to-4.6 gain was 5 points, and vendor pre-release claims routinely overshoot independent runs.

Why: Grok moved from 56 to 61 on the index across a full version bump, so a single point-release beating both the current co-leader (Sol at 61) and Opus 5 (which sits above both) means clearing several points in one 3-to-4-week cycle. This episode already gives the pattern for how vendor claims land against neutral harnesses: DeepSeek V4 Pro's leaked 87.9% on TerminalBench collapsed to a 53 on the independent index. Musk's "SpaceX training corpus" pitch is a capability claim with no eval behind it, and the burden is on the model to clear the top of a field where gains have gotten expensive. The opposite outcome, 4.7 genuinely topping everything, would require the largest one-release jump anyone's posted this cycle, which is the less likely bet.

Revisit by 2026-09-18: We're right if Grok 4.7's independent Artificial Analysis index score fails to exceed both GPT-5.6 Sol and Claude Opus 5. We're wrong if the independent run puts 4.7 clearly ahead of both.

Comments