Industry story
Open-Source AI Models Halving Gap-Closure Time Each Era
agents evals inference model-pricing open-weights
A SemiAnalysis analysis spanning three distinct LLM eras — early scaling (2022–2024), reasoning (2024–2025), and agentic (2025–present) — finds that open-source models are closing the capability gap with frontier closed models at an accelerating rate. Specifically, with each new era, open models take roughly half as long to match the leading closed model: ~17 months in Era 1, ~8.5 months in Era 2, and under 6 months in Era 3. The authors benchmarked models using era-appropriate evaluation suites (e.g., GSM8K/MMLU-Pro for Era 1, AIME/HLE for Era 2, SWE-bench/BrowseComp for Era 3) to avoid benchmark saturation artifacts.
In the current agentic era, Kimi K2.6 surpassed Anthropic's Opus 4.5 in 4.8 months and Zhipu's GLM-5.2 cleared OpenAI's GPT-5.2 in 6 months on the authors' composite benchmark suite. The authors note important caveats: benchmarks can be hill-climbed via RL (reinforcement learning) environments that mimic test tasks, and real-world usability (e.g., Claude Code's harness) still differentiates closed models even when open models score comparably. Nevertheless, the consistent halving trend raises questions about long-term monetization viability for frontier labs if open models remain competitive at far lower cost.
Full analysis
SemiAnalysis has a chart that says open models are catching up to the frontier faster every era: 17 months, then 8.5, then under 6. Kimi K2.6 passed Opus 4.5 in 4.8 months. GLM-5.2 passed GPT-5.2 in 6. The question for anyone building on models is whether that curve is a law you can plan around, or a two-points-per-era line that only looks like one.
The Skeptic. Three eras, two observations each, and we're calling it a geometric trend. That's six data points fit to a curve everyone wanted to see. For the halving to hold, closed labs have to keep dropping one tidy "era-defining" model at fixed intervals, open development has to stay unconstrained by compute, and benchmarks can't get gamed faster than they get retired. None of those is clean. Dylan Patel's own team flags the RL hill-climbing problem: you can train a model to look like the test. And Kimi and Zhipu aren't garage projects. They're running enormous RL clusters on state-adjacent capital. For the PM in the room: the models tie on the exam, not on the job.
The Compute Pragmatist. The framing undersells that this is a compute story. Closing the gap in five months takes big RL runs, not clever fine-tuning. DeepSeek's recipe got copied, and copying it needs GPUs. So "open" no longer means "cheap to make." It means a different spend pattern: less pretraining FLOP, more RL, faster loops. That compresses the window where an H200 cluster earns a closed lab differentiated return. If Era 4 closes in under three months, the training-compute moat expires before the next big run even finishes. For the PM: the moat was time, and time is getting cheap.
The Builder. Benchmark parity is not integration parity, and that gap is where your on-call engineer lives. Swap an open model that "matches" Opus into an agent loop and you inherit worse tool-call reliability, weaker coherence over long trajectories, and refusal behavior nobody tuned. Claude Code's harness is the thing SemiAnalysis admits still separates the leaders. The eval says tie; production says budget 90 days of regression. The upside is real: near-frontier weights can cut an inference bill 70 to 90%. But the savings are visible on day one and the silent failures show up in week six. For the PM: the exam score is free, the deployment isn't.
The Enterprise Buyer. A CTO doesn't sign a contract because a Chinese open model cleared GPT-5.2 on a composite suite. They sign for indemnification, audit logs, data residency, and someone to call when the agent deletes a production table. Kimi K2.6 and GLM-5.2 carry a second problem in most Western procurement: provenance. Open weights mean you own the alignment posture, which enterprises read as "you own the liability." The halving curve tells buyers the capability floor is rising everywhere. It does not tell them the vendor risk is falling. For the PM: the model got good; the paperwork got harder.
Where they split
Two real disagreements. The Compute Pragmatist says the moat is expiring on a clock and closed labs should be updating their thesis now. The Enterprise Buyer says capability parity barely moves the purchase, because what enterprises buy is risk transfer, and open weights make that liability position worse. Both can be true: the technical moat collapses while the commercial moat holds. That's the actual state of play.
The second split is the Skeptic against the chart itself. Is "capability gap" even the same measurement across three eras with different task structures? You're comparing MMLU-era scaling to agentic SWE-bench trajectories and calling both "the gap." The trend is only as real as the construct, and the construct is shakier than the curve looks.
What it hinges on
The decision for a builder isn't "believe the halving or don't." It's narrower: does open-model parity on your workload arrive before your closed-model contract renews, and does it arrive with enough harness quality that you don't eat a quarter of reliability debt? That splits into two testable things. First, whether the parity claim holds on your task distribution, not AIME or BrowseComp. Second, whether the production harness gap closes as fast as the benchmark gap, which it has not so far.
The council leans one way: the capability floor is genuinely rising fast, and the cost argument for open weights on batch and non-latency-critical agent work is already strong enough to pilot. But the extrapolation to "closed labs are cooked" is the Skeptic's overfit trap, and the Enterprise Buyer's point kills the clean version of it.
Before you commit: run your own eval on your own agent trajectories, not the paper's suite. Load-test tool-call reliability at production concurrency, because that's the failure that doesn't show in a leaderboard. And if it's a Chinese open model, get legal on provenance before eng gets attached.
The Prediction
Prediction: On the next SemiAnalysis update to this analysis (or an equivalent third-party agentic benchmark refresh) before 2027-02-28, the top open-weight model will match or beat the leading closed model on composite agentic scores, but no independent production-harness test (tool-call reliability, long-trajectory coherence at load) will show that open model within 10% of the closed leader.
Confidence: Medium. Score parity is on-trend; harness parity has never followed it.
Why: The halving curve is measuring benchmark scores, and RL hill-climbing plus fast open iteration make continued score convergence the likely outcome; that part I'd bet with the trend. But SemiAnalysis itself concedes the differentiator that survives is the harness, Claude Code being the named example, and harness quality comes from integration engineering and reliability tuning that open weights don't ship with. The score gap and the production gap are closing at different rates, and nothing in this data suggests the harness gap is halving too. The opposite outcome, an open model matching the closed leader on real agent reliability at load, would require a jump nobody has demonstrated yet.
Revisit by 2027-02-28: We're right if the next agentic benchmark cycle shows an open model at or above the closed leader on scores while independent harness/reliability testing still shows a double-digit gap. We're wrong if an open-weight model demonstrably matches the closed leader on production tool-call reliability and long-trajectory coherence, not just benchmark scores.
The score race and the reliability race are different races. The chart tracks the one that's easier to win.
Comments