Industry story
Open-Source AI Models Halving Gap-Closure Time Each Era
agents evals inference model-pricing open-weights
A SemiAnalysis analysis spanning three distinct LLM eras — early scaling (2022–2024), reasoning (2024–2025), and agentic (2025–present) — finds that open-source models are closing the capability gap with frontier closed models at an accelerating rate. Specifically, with each new era, open models take roughly half as long to match the leading closed model: ~17 months in Era 1, ~8.5 months in Era 2, and under 6 months in Era 3. The authors benchmarked models using era-appropriate evaluation suites (e.g., GSM8K/MMLU-Pro for Era 1, AIME/HLE for Era 2, SWE-bench/BrowseComp for Era 3) to avoid benchmark saturation artifacts.
In the current agentic era, Kimi K2.6 surpassed Anthropic's Opus 4.5 in 4.8 months and Zhipu's GLM-5.2 cleared OpenAI's GPT-5.2 in 6 months on the authors' composite benchmark suite. The authors note important caveats: benchmarks can be hill-climbed via RL (reinforcement learning) environments that mimic test tasks, and real-world usability (e.g., Claude Code's harness) still differentiates closed models even when open models score comparably. Nevertheless, the consistent halving trend raises questions about long-term monetization viability for frontier labs if open models remain competitive at far lower cost.
Analysis
Showing the shorter version.
SemiAnalysis has a chart showing open models closing the gap to frontier models faster each era: 17 months, then 8.5, then under 6. Kimi K2.6 passed Opus 4.5 in 4.8 months. GLM-5.2 passed GPT-5.2 in 6. The question is whether that compression is a plannable curve or a three-era pattern that someone decided looks like a law.
Two reasons to be skeptical of the extrapolation. First, the trend assumes closed labs keep dropping tidy "era-defining" models at fixed intervals, open development stays unconstrained by compute, and benchmarks don't get gamed faster than they're retired. None of those hold cleanly. Dylan Patel's own team flags the RL hill-climbing problem: you can train a model to score well without building the underlying capability. Second, Kimi and Zhipu aren't garage projects. They're running enormous RL clusters on state-adjacent capital. "Open" no longer means cheap to make. It means a different spend pattern: less pretraining compute, more RL, faster loops.
That compute shift matters for the closed-lab moat. If Era 4 closes in under three months, the training-compute advantage expires before the next big pretraining run even finishes. The moat was time, and time is compressing.
But benchmark parity and production parity are different things, and that gap is where your on-call engineer lives. Swap an open model that "matches" Opus into an agent loop and you inherit worse tool-call reliability, weaker coherence over long trajectories, and refusal behavior nobody tuned. SemiAnalysis itself points to Claude Code's harness as what still separates the leaders. The eval says tie; production says budget 90 days of regression. The inference cost savings are real, 70 to 90% on an open-weight model, but they're visible on day one and the silent failures show up in week six.
Enterprise procurement adds another layer. A CTO doesn't sign because a Chinese open model cleared GPT-5.2 on a composite benchmark suite. They sign for indemnification, audit logs, data residency, and someone to call when the agent deletes a production table. Open weights mean you own the alignment posture, which procurement reads as owning the liability. The capability floor rising doesn't make the vendor risk fall.
Both things can be true simultaneously: the technical moat collapses while the commercial moat holds. That's the actual state of play right now.
The call: On the next major agentic benchmark refresh before 2027-02-28, the top open-weight model will match or beat the leading closed model on composite agentic scores. No independent production-harness test, covering tool-call reliability and long-trajectory coherence under load, will show the open model within 10% of the closed leader. Medium confidence. Score parity is on-trend; harness parity has never followed it, and nothing in this data suggests it's halving too.
If you're building on models, the practical question isn't whether to believe the halving curve. It's narrower: does open-model parity on your specific workload arrive before your closed-model contract renews, and does it arrive with enough harness quality that you don't eat a quarter of reliability debt? Run your own eval on your own agent trajectories, not the paper's suite. Load-test tool-call reliability at production concurrency. And if it's a Chinese open model, get legal on provenance before engineering gets attached.
The score race and the reliability race are different races. The chart tracks the one that's easier to win.
Your draft
SemiAnalysis has a chart that says open models are catching up to the frontier faster every era: 17 months, then 8.5, then under 6. Kimi K2.6 passed Opus 4.5 in 4.8 months. GLM-5.2 passed GPT-5.2 in 6. The question for anyone building on models is whether that curve is a law you can plan around, or a two-points-per-era line that only looks like one.
The Skeptic. Three eras, two observations each, and we're calling it a geometric trend. That's six data points fit to a curve everyone wanted to see. For the halving to hold, closed labs have to keep dropping one tidy "era-defining" model at fixed intervals, open development has to stay unconstrained by compute, and benchmarks can't get gamed faster than they get retired. None of those is clean. Dylan Patel's own team flags the RL hill-climbing problem: you can train a model to look like the test. And Kimi and Zhipu aren't garage projects. They're running enormous RL clusters on state-adjacent capital. For the PM in the room: the models tie on the exam, not on the job.
The Compute Pragmatist. The framing undersells that this is a compute story. Closing the gap in five months takes big RL runs, not clever fine-tuning. DeepSeek's recipe got copied, and copying it needs GPUs. So "open" no longer means "cheap to make." It means a different spend pattern: less pretraining FLOP, more RL, faster loops. That compresses the window where an H200 cluster earns a closed lab differentiated return. If Era 4 closes in under three months, the training-compute moat expires before the next big run even finishes. For the PM: the moat was time, and time is getting cheap.
The Builder. Benchmark parity is not integration parity, and that gap is where your on-call engineer lives. Swap an open model that "matches" Opus into an agent loop and you inherit worse tool-call reliability, weaker coherence over long trajectories, and refusal behavior nobody tuned. Claude Code's harness is the thing SemiAnalysis admits still separates the leaders. The eval says tie; production says budget 90 days of regression. The upside is real: near-frontier weights can cut an inference bill 70 to 90%. But the savings are visible on day one and the silent failures show up in week six. For the PM: the exam score is free, the deployment isn't.
The Enterprise Buyer. A CTO doesn't sign a contract because a Chinese open model cleared GPT-5.2 on a composite suite. They sign for indemnification, audit logs, data residency, and someone to call when the agent deletes a production table. Kimi K2.6 and GLM-5.2 carry a second problem in most Western procurement: provenance. Open weights mean you own the alignment posture, which enterprises read as "you own the liability." The halving curve tells buyers the capability floor is rising everywhere. It does not tell them the vendor risk is falling. For the PM: the model got good; the paperwork got harder.
Where they split
Two real disagreements. The Compute Pragmatist says the moat is expiring on a clock and closed labs should be updating their thesis now. The Enterprise Buyer says capability parity barely moves the purchase, because what enterprises buy is risk transfer, and open weights make that liability position worse. Both can be true: the technical moat collapses while the commercial moat holds. That's the actual state of play.
The second split is the Skeptic against the chart itself. Is "capability gap" even the same measurement across three eras with different task structures? You're comparing MMLU-era scaling to agentic SWE-bench trajectories and calling both "the gap." The trend is only as real as the construct, and the construct is shakier than the curve looks.
What it hinges on
The decision for a builder isn't "believe the halving or don't." It's narrower: does open-model parity on your workload arrive before your closed-model contract renews, and does it arrive with enough harness quality that you don't eat a quarter of reliability debt? That splits into two testable things. First, whether the parity claim holds on your task distribution, not AIME or BrowseComp. Second, whether the production harness gap closes as fast as the benchmark gap, which it has not so far.
The council leans one way: the capability floor is genuinely rising fast, and the cost argument for open weights on batch and non-latency-critical agent work is already strong enough to pilot. But the extrapolation to "closed labs are cooked" is the Skeptic's overfit trap, and the Enterprise Buyer's point kills the clean version of it.
Before you commit: run your own eval on your own agent trajectories, not the paper's suite. Load-test tool-call reliability at production concurrency, because that's the failure that doesn't show in a leaderboard. And if it's a Chinese open model, get legal on provenance before eng gets attached.
The Prediction
Prediction: On the next SemiAnalysis update to this analysis (or an equivalent third-party agentic benchmark refresh) before 2027-02-28, the top open-weight model will match or beat the leading closed model on composite agentic scores, but no independent production-harness test (tool-call reliability, long-trajectory coherence at load) will show that open model within 10% of the closed leader.
Confidence: Medium. Score parity is on-trend; harness parity has never followed it.
Why: The halving curve is measuring benchmark scores, and RL hill-climbing plus fast open iteration make continued score convergence the likely outcome; that part I'd bet with the trend. But SemiAnalysis itself concedes the differentiator that survives is the harness, Claude Code being the named example, and harness quality comes from integration engineering and reliability tuning that open weights don't ship with. The score gap and the production gap are closing at different rates, and nothing in this data suggests the harness gap is halving too. The opposite outcome, an open model matching the closed leader on real agent reliability at load, would require a jump nobody has demonstrated yet.
Revisit by 2027-02-28: We're right if the next agentic benchmark cycle shows an open model at or above the closed leader on scores while independent harness/reliability testing still shows a double-digit gap. We're wrong if an open-weight model demonstrably matches the closed leader on production tool-call reliability and long-trajectory coherence, not just benchmark scores.
The score race and the reliability race are different races. The chart tracks the one that's easier to win.
Also covered this issue
-
Data center opposition surges 33 points in a year, Senate Republicans warn of political blowback
transformer-news
US data center opposition is hardening into a political cost that could delay or kill your compute infrastructure timeline by years.
-
SemiAnalysis Launches AgentX 1.0: First Open-Source Agentic Inference Benchmark
semianalysis
Agentic workloads are now burning 10–100x more tokens per task, forcing you to re-model inference costs and hardware allocation before your next contract renewal.
Comments