Podcast episode
What Chess.com Teaches US About Superhuman Capabilities, with CEO Erik Allebest
agents coding-agents evals open-weights
Erik Allebest, CEO of Chess.com, joins Sarah Guo and Elad Gil to talk through what 25 years of superhuman chess engines actually did to the game. The short answer: nothing bad. Chess grew to 250 million registered players and $200M in revenue after machines left humans in the dust.
Allebest's most transferable observation isn't the growth numbers, which had help from COVID and The Queen's Gambit. It's the Leela Chess Zero story: when neural-net engines replaced hand-tuned deterministic ones like Stockfish, top play got weirder and more watchable. The most accurate engine produced the most boring chess. That tension between optimal and engaging is real for anyone tuning a recommender or bidding model that users actually interact with. Allebest also claims agentic coding (AI that writes and tests its own code) is already compressing their spec-to-ship cycle at Chess.com's 650-person shop, though he offers no actual numbers.
The chess-grew-therefore-your-product-is-safe argument doesn't hold. Chess has a scoreboard; people play for pride. Most software doesn't. Steal the accuracy-versus-engagement lesson, note that agentic coding is now table stakes, then go measure your own cycle time.
Full analysis
Your draft
Erik Allebest ran a 20-year experiment nobody designed on purpose. Machines went superhuman at chess in 1997, and instead of killing the game, they grew it to 250 million registered players and $200M in revenue. His claim is that the same thing happens everywhere AI beats us: humans keep wanting to do human stuff, and skill stays valuable. That's the whole story worth pulling out for anyone shipping AI. The rest is a founder profile.
What's actually being decided here for a build team: nothing forcing. This is a Type 2 read. No model release, no pricing change, no deprecation. It's a data point you either update on or ignore. The forcing function is only your own planning cycle, whenever you next argue about whether "AI will make our product obsolete" deserves airtime.
The Skeptic. One firm's engagement chart is not a law of nature. Chess had a moat AI can't cross: humans want to beat other humans at a bounded game with a scoreboard. That's not most software. Allebest's own COVID and Queen's Gambit bumps did more for the numbers than any thesis about superhuman AI, and he admits both left a higher baseline. For a PM: chess grew because people play each other for fun, not because the engine got smart. Don't generalize "engagement went up" to your CRUD app or your ad-tech pipeline where nobody is playing for pride.
The Researcher. The genuinely interesting technical note is the Leela Chess Zero story. When Stockfish, a hand-tuned deterministic engine, dominated, top play got "boring" because the answers felt mechanical. Neural-net engines that learned from self-play revived interest because they played weird, human-surprising moves. For a builder: the same brute-force optimum can be correct and unwatchable, while a learned model that's slightly worse but more varied drives more engagement. That maps onto a real product tension. Your most accurate recommender or most optimal bid-shading model may be the one users find least interesting to interact with.
The Builder. The one operational claim I can use: Allebest says agentic development, where the AI writes and tests code on its own, is already compressing their spec-to-ship cycle at a 650-person, fully-remote company. No frontier gear, no named lab, just applied tooling. That's the adoption signal. Agentic coding has left the demo and is doing real work at mid-scale shops. The caution: he gives no numbers. "Great success shortening" the cycle is a vibe, not a metric. Before you tell your team this validates your own agentic push, ask him or yourself what the actual cycle-time delta is and where the QA still catches the AI's mistakes.
The Open-Source Advocate. Buried in the chess history is a cleaner point than Allebest makes. Leela Chess Zero is an open, community-run reimplementation of DeepMind's AlphaZero approach. A volunteer project reproduced a frontier lab's headline result and now sits at the top of the game. That's the reproducibility story worth carrying into any AI-strategy meeting. When the method is public, the closed lab's lead in a bounded domain compresses fast. For an ad-tech team weighing build-versus-buy on models, it's a reminder that today's proprietary edge is often next year's open weight.
Where these part ways: The Skeptic says the chess result doesn't transfer, and he's mostly right about the mechanism. The Researcher says one thing does transfer, the finding that optimal isn't the same as engaging, and that one's real for anyone tuning a model users actually touch. The Builder wants the agentic-dev claim to mean something, but without a cycle-time number it's a testimonial, not evidence.
What this actually hinges on: whether you treat Chess.com as proof that AI grows human markets, or as a single domain with a scoreboard that AI happens not to have ruined. The second reading is the right one. Don't put "chess grew after Deep Blue" in a strategy deck as if it forecasts your product. Do steal the Leela lesson about accuracy versus interestingness, and do note that agentic coding is now table stakes at companies your size, then go measure your own ship-cycle delta instead of trusting his.
Impact on your roadmap is low. This is a good listen for framing, not a signal that changes what you build next quarter.
No high-conviction prediction this week.
The episode is a founder profile with a few applied-AI asides. Nothing here is a dated, falsifiable claim about model capability or the AI market that a build team could check in three months. Allebest gives no timeline on AGI, no numbers on his agentic-dev gains, and no product ship dates. Forcing a prediction out of this would mean betting on something the source doesn't actually cover.
Comments