Podcast episode
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
agents evals inference model-pricing
Joon Sung Park, the Stanford researcher behind the 2023 "Smallville" generative-agents paper, is on Latent Space to pitch Simile AI: foundation models trained to predict how real, irrational humans behave, not to be helpful or correct. The core claim is a benchmark gap. Frontier models like GPT and Claude score 20 to 60% at predicting the behavior of specific populations; Simile says it hits 85%, trained on life-story interviews, transaction data, and randomized controlled trials (real experiments that measure why people choose, not just what they say). They just raised $2 billion.
The skepticism writes itself. "85% of what, exactly?" is the right question. Park's benchmark measures whether the AI reproduces how participants answered questions two weeks after the original interview. That ceiling is human self-consistency, which is noisy. And the 20% frontier-model baseline is a task those models were never trained for, so the gap is partly a training-objective artifact, not proof behavior needs a dedicated $2B model.
The part worth filing away isn't the benchmark. It's the workload: massively parallel, latency-insensitive batch inference. If behavioral simulation becomes a real product category, the economics favor whoever has idle cheap compute, not whoever owns the fanciest cluster.
Analysis
Showing the shorter version.
Simulation as a Scaling Law
Joon Sung Park, the Stanford researcher behind the 2023 "Smallville" generative-agents paper, is now running Simile AI. The pitch: foundation models trained to predict how real, irrational humans behave, not to be helpful or correct. They just raised $2B on that thesis, and the headline number is an 85% score on behavioral prediction against frontier models that score 20 to 60% on the same task.
Before you file that away as impressive, ask what 85% is measuring. Simile's benchmark grades digital twins against how well real participants reproduce their own answers two weeks later. Human test-retest consistency is already noisy, so 85% of a wobbly ceiling is a smaller number than it sounds. And "frontier models score 20%" is a well-worn move: benchmark an incumbent on a task it was never trained for and call the gap a product. RLHF-tuned models are sycophantic and mode-collapsed by design; of course they predict niche populations badly.
The training data is what's genuinely new here. Post-training on pre-registered randomized controlled trials (RCTs) from the Open Science Framework carries causal signal: why people chose, not just what they said. Frontier labs are not systematically ingesting that. If Park's claim holds, that behavioral scaling yields predictable gains on a separate curve from language-model loss, that is a real differentiable axis. But the episode never shows a held-out prediction of a future RCT the model never saw. Reproducing published PNAS results after training on the same public corpus that contains those results is closer to memorization than prediction. That is the question the whole thing hinges on.
The moat question also splits. Simile's real asset is probably the consent-based two-hour life-story interviews with 1,000 US-representative people, something you cannot scrape off Hugging Face. The OSF corpus itself is public, so the technique is reproducible; the interview layer is not. If the moat is the interview data, Simile has a defensible position. If the durable asset is whoever can run society-scale batch inference cheapest, any hyperscaler with idle GPUs eats this inside eighteen months.
For operators, the inference workload is worth noting. Tens of millions of parallel agent simulations are embarrassingly parallel, batch, and latency-insensitive. The economics point toward cheap off-peak commodity GPU capacity, not premium low-latency H100 time. That is the opposite profile from real-time bidding.
Nothing from Simile is shippable Tuesday; it is a closed enterprise service, not an API. But the idea is testable. If your team generates synthetic users for market research or cold-start personalization, you already know GPT returns the bland median human every time. Take an open base model, post-train on a slice of OSF RCTs, and check whether your synthetic panel stops collapsing to the median. That experiment costs a weekend.
The call: No independent replication published before June 30, 2027 will show Simile's behavioral model beating a frontier model by more than 15 points on a held-out behavioral-prediction task using RCTs the model was not trained on. Medium confidence. Skeptics have every reason to run that falsifying experiment; a company that just raised $2B on the current framing does not.
Your draft
Joon Sung Park, the Stanford researcher behind the 2023 "Smallville" generative-agents paper, is on Latent Space to pitch Simile AI: foundation models trained to predict how real, irrational humans behave, not to be helpful or correct. The claim that should make an operator sit up is a benchmark gap. Frontier models (GPT, Claude) score 20 to 60% at predicting the behavior of niche populations; Simile says it hits 85% by post-training on interviews, transaction data, and randomized controlled trials. And they just raised $2B to do it.
Reversibility for a builder: Type 2 for now. Nobody is being asked to bet an architecture on this. The question is whether "behavioral simulation" is a real new axis of model capability you should start evaluating, or a well-funded research demo wearing an enterprise logo. There is no forcing function. This is a "should I run my own test against this claim" decision, not a procurement one.
The Skeptic
Eighty-five percent of what, exactly? The 1,000-people paper measures digital twins against how well participants reproduce their own answers two weeks later. So the ceiling is human test-retest noise, and 85% of a noisy ceiling is a smaller number than it sounds. For a PM: they graded the AI against a wobbly human baseline, not against truth. And "frontier models score 20%" is the oldest move in the book, benchmark your incumbent on a task it was never trained for, then declare a gap. RLHF-tuned models are sycophantic and mode-collapsed; of course they predict the median human badly. That is a training-objective artifact, not proof that behavior needs a $2B model. The UBI tell is buried in the episode: Sam Altman funded a $40M, five-year real-world UBI study that concluded it did not work as hoped. If a real RCT with real humans surprised its funder, why trust a simulator to have predicted it?
The Researcher
The genuinely new thing here is the training objective. Post-training on pre-registered RCTs from the Open Science Framework is a distinct signal from RLHF: RCTs carry causal labels, why people chose, not just what they said. That is a real data type frontier labs are not systematically ingesting. The scaling-law claim is the interesting research bet: Park says more human data plus compute yields predictable simulation gains, a separate curve from language-model loss. If that replicates, it is a differentiable axis. But note what is missing. No held-out prediction of a future RCT the model never saw. Reproducing published PNAS results after training on the OSF corpus risks contamination, if those studies are in the training set, you are testing memorization dressed as prediction. The whole edifice hinges on out-of-sample causal prediction, and the episode does not show it.
The Open-Source Advocate
Here is the part nobody can reproduce: the data. The model architecture is almost beside the point. What Simile actually owns is consent-based two-hour life-story interviews with 1,000 US-representative people, plus curated RCTs. You cannot pull that off Hugging Face. That is the moat, and it is a moat made of recruiting and IRB paperwork, not weights. For the open ecosystem this cuts two ways. The OSF corpus itself is public, so anyone can try RCT post-training on an open base model (Llama, Qwen, Mistral) and see if the behavioral-prediction lift shows up. That experiment is cheap and worth running. But the interview layer is proprietary and expensive, so the technique is reproducible and the product is not. Expect a paper within a year showing open weights plus OSF post-training closes some of the gap for cheap.
The Compute Pragmatist
Park is describing a new workload class, and this is the part operators should actually file away. Not training, inference. "Tens of millions of simulations," and a stated ambition to run society-scale simulations that cost as much as training a frontier model. That is massively parallel multi-agent inference: embarrassingly parallel, batch, latency-insensitive. For a PM: it is the opposite of real-time bidding, you do not care if an answer takes an hour, you care about throughput per dollar. That maps beautifully onto cheap, off-peak, commodity GPU capacity and spot markets, not premium low-latency H100 time. If behavioral simulation becomes a real product category, the economics favor whoever has idle batch compute, which is every hyperscaler with a trough to fill. The $2B is a bet that enterprises and governments will pay training-run prices for a decision-support answer. That is a big if.
The Builder
What would I ship Tuesday? Nothing from Simile, it is a closed enterprise service with Fortune 100 logos, not an API you drop into a sprint. But the idea is stealable. If your team does anything with synthetic users, market-research automation, or cold-start personalization, you are already using GPT to fake a user, and it gives you the bland median human every time. The actionable move is small: take an open base model, post-train on a slice of OSF RCTs, and check whether your synthetic-user panel gets less mode-collapsed on your own task. Cheap experiment, clear read. Do not wait for Simile to sell you anything.
Where they disagree
The Researcher sees a real new training axis; the Skeptic sees a benchmark chosen to flatter it. That fight resolves on one thing: out-of-sample causal prediction. Reproducing studies that may be in the training data proves nothing.
The Open-Source Advocate and the Compute Pragmatist split on where the moat lives. One says it is the un-scrapeable interview data (durable). The other says the durable asset is whoever can run society-scale batch inference cheapest (contestable by any hyperscaler). If the moat is data, Simile is defensible. If it is compute, a frontier lab with the OSF corpus and a spare cluster eats this in eighteen months.
What it hinges on
Does RCT post-training produce out-of-sample behavioral-prediction gains on studies the model never saw? If yes, this is a real capability axis and the $2B is early, not crazy. If the 85% is memorization of a public corpus plus a soft human-retest ceiling, it is a very expensive research demo. The council leans skeptical on the product and curious on the technique. The technique is cheap enough to test yourself, which is the whole point.
Prediction: No independent replication published before 2027-06-30 will show Simile AI's behavioral foundation model beating a frontier model (GPT, Claude, Gemini) by more than 15 points on a held-out behavioral-prediction task using RCTs the model was not trained on.
Confidence: Medium. The technique is plausible, but the published wins rely on possibly-contaminated public data.
Why: Simile's headline gap comes from reproducing PNAS and General Social Survey results after post-training on the Open Science Framework, a public corpus those very results may live in, so the demonstrated lift could be memorization rather than prediction. The claim that survives scrutiny is out-of-sample causal prediction, and nothing in the episode shows it; Park frames the 85% against human test-retest fidelity, a noisy ceiling that inflates the number. A truly held-out test, predicting an RCT outcome the model never ingested, is the one experiment that would settle it, and closed enterprise vendors rarely run the experiment that could disprove their pitch. The opposite outcome, a clean third-party held-out win, is possible if the scaling-law claim is real, but skeptics have the incentive to publish that test, not a company that just raised $2B on the current framing.
Revisit by 2027-06-30: We're right if no peer-reviewed or preprint replication shows a >15-point held-out advantage over a frontier model on unseen behavioral-prediction tasks. We're wrong if such a replication appears and holds up.
The cheap tell for any operator in the meantime: post-train an open model on a slice of OSF and check whether your own synthetic users stop collapsing to the median. That experiment costs a weekend and tells you more than the $2B headline does.
Also covered this issue
-
Open-Source AI Models Halving Gap-Closure Time Each Era
semianalysis
Open models are closing benchmark gaps in months, but production reliability gaps stay wide, making eval parity a trap for teams planning migrations before reliability catches up.
-
Data center opposition surges 33 points in a year, Senate Republicans warn of political blowback
transformer-news
US data center opposition is hardening into a political cost that could delay or kill your compute infrastructure timeline by years.
-
SemiAnalysis Launches AgentX 1.0: First Open-Source Agentic Inference Benchmark
semianalysis
Agentic workloads are now burning 10–100x more tokens per task, forcing you to re-model inference costs and hardware allocation before your next contract renewal.
Comments