Industry story
UK AISI and EvalEval release standardized benchmark results for frontier models
evals inference measurement model-pricing security
The UK AI Security Institute (AISI) and the EvalEval Coalition have published verified evaluation results for five major AI benchmarks—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0—covering six frontier models including Claude Opus 4/4.5/4.6, GPT-5, GPT-5.2, and GPT-5.4. The release accompanies AISI's paper 'How Inference Compute Shapes Frontier LLM Evaluation,' which examines how benchmark performance shifts depending on inference-time compute (the amount of computation a model uses at the moment it generates an answer) and evaluation protocol choices.
The initiative addresses a reproducibility crisis in AI evaluation: results are currently reported across many formats without enough configuration detail to replicate them, and re-running evals can be prohibitively expensive. AISI and EvalEval are jointly building shared infrastructure—including the 'Every Eval Ever' (EEE) reporting schema and an 'Evaluation Cards' platform—to bring results and their interpretive context into a common structure, enabling reliable cross-study comparisons and meta-research relevant to both AI governance and deployment decisions.
Full analysis
The UK's AI Security Institute and the EvalEval Coalition published verified scores for six frontier models across five hard benchmarks, plus a paper showing that a model's benchmark rank moves depending on how much computation it burns at answer time. The pitch is a shared schema so anyone can reproduce a benchmark claim. What's actually being decided is whether the numbers you use to pick a model come with the fine print that makes them mean anything.
This is easy to undo for you. Nobody has to adopt anything. That's also the whole problem. There is no deadline and no enforcement. The labs submit their own numbers and control the inference budget those numbers were run at.
The Skeptic
Five benchmarks, six models, one paper. That's a coordination announcement that treats a solved problem as solved when it isn't. AISI has no enforcement. EvalEval is a coalition, not a regulator. The labs still control the weights and the inference budgets, so every "standardized" score is a vendor-submitted number that gets audited after the fact, if at all. HealthBench and Humanity's Last Exam already smell of contamination, where labs train on data close enough to the test that the score inflates. And the admission that re-running these evals is "prohibitively expensive" tells you who verifies: nobody without a lab-sized budget. Good-faith infrastructure a sophisticated actor routes around in an afternoon.
The Compute Pragmatist
The paper is the part that pays rent. Frontier models now win benchmarks by spending at answer time: longer reasoning chains, repeated sampling, verifier loops. So a model that tops Humanity's Last Exam at 100,000 reasoning tokens can be third-best at the token budget your production system actually runs. AISI quantifying that link hands buyers something real: a way to re-rank scores against the compute envelope you'll actually pay for. Peak score and deployment cost have quietly divorced. Claude Opus 4.6's headline number and Claude Opus 4.6's number at your budget are different animals, and until now the leaderboard hid the difference.
The Researcher
The reporting schema is the correct abstraction. The field has drowned in benchmark claims nobody can rerun for two years. What's underrated is the inference-compute confound: the same weights under extended chain-of-thought at ten times the token budget look like a categorically better model, and almost no public leaderboard controls for it. Making that variable explicit is the methodological win. The prize isn't any single score. It's the meta-research that becomes possible once results share a structure: which benchmarks correlate, how much elicitation effort swings a number, where a "capability" is really a compute artifact.
The Enterprise Buyer
Here's the gap between the announcement and the contract. Even verified, comparable scores don't tell a CTO what they need. I'm not buying a HealthBench number. I'm buying uptime, data handling, and a price per million tokens I can forecast. What this release does give me is leverage in the room. When a vendor waves a peak benchmark score, I can now ask for the score at my inference budget and cite AISI's own paper when they dodge. That's a negotiating input. But adoption is voluntary, so the labs most confident in their peak numbers are the ones least motivated to publish the compute-adjusted version that makes them look ordinary.
The Safety Lens
The compute-sensitivity finding is a live safety issue that the footnotes have been obscuring. A model that refuses to help with dangerous chemistry at standard decoding can behave differently when you let it reason longer. That gap is exactly where misuse lives, and governance frameworks have been scoring the wrong configuration. Making compute a first-class variable in the report closes a loophole red-teamers already knew about. The Evaluation Cards platform, if labs adopt it, builds the audit trail that pre-deployment safety cases need. The weak link is the same one everywhere in this story: voluntary participation. The most capable future models are the ones with the most to lose from a full evaluation card, and nothing forces them onto it.
Where they disagree
The real split is between the Researcher and the Skeptic, and it's about whether a schema without teeth changes behavior. The Researcher sees a genuine methodological unlock: make compute explicit and a whole layer of junk claims becomes visible. The Skeptic says visibility is worthless when the actor being measured supplies the number and can decline to play. Both are right, which tells you the outcome hinges on adoption, not on the quality of the standard.
The second tension is the Compute Pragmatist against the Enterprise Buyer. The Pragmatist says compute-adjusted scores are a purchasing input right now. The Buyer says a score is never the thing being bought, and the labs with the best peak numbers have the least reason to publish the adjusted ones. The compute-adjusted picture is genuinely useful and structurally under-supplied at the same time.
What it hinges on
One belief: will the labs whose peak scores look best under generous compute voluntarily publish the compute-adjusted numbers that make them look average? Everything else in the release is sound. The schema is good, the paper is good, the safety logic is good. None of it binds anyone. If you're using these benchmarks to pick a model, do the thing the labs won't do for you: rerun the two evals that matter to your use case at the token budget your production system actually runs, and rank on that. The AISI paper hands you the reason. The number you get will not match the leaderboard, and that mismatch is the whole point.
Prediction: By the time AISI and EvalEval publish their next round of verified frontier-model results after this September 2026 release, at least one of OpenAI or Anthropic will still report a headline benchmark score for a flagship model (GPT-5.x or Claude Opus 4.x class) without a matching compute-adjusted figure in the EEE schema.
Confidence: Medium. Participation is voluntary and peak scores sell better than adjusted ones.
Why: The release itself admits reproducing these evals is "prohibitively expensive" and that AISI has no enforcement, so publishing the compute-adjusted number is a choice, not a requirement. A lab's headline number is a marketing asset, and the AISI paper shows that same number shrinks when you normalize to a realistic inference budget, so voluntarily publishing the adjusted figure means voluntarily deflating your own leaderboard rank right as competitive stakes on those rankings are rising. The opposite outcome, both labs fully populating the schema with compute-adjusted scores on their flagships, would require them to hand rivals and buyers a cleaner way to discount their best marketing numbers, with nothing forcing them to. Labs already report peak scores in press materials and leave the fine print to appendices; that habit is the path of least resistance here too.
Revisit by 2027-03-28: We're right if the next AISI/EvalEval verified results round shows any flagship GPT-5.x or Claude Opus 4.x model with a headline score but no compute-adjusted figure in the shared schema. We're wrong if both OpenAI and Anthropic supply compute-adjusted numbers for every flagship model they submit.
Comments