Refacto AI

Podcast episode

🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery

diffusion-models gpu-supply open-weights protein-design structural-biology

Latent Space hosts RJ Honicky and Joshua Meier sit down with Chai Discovery co-founders Matthew McPartlon and Neil Patil to talk about what it actually means to design antibodies with AI, not just predict how proteins fold.

McPartlon and Patil's case is that structural biology AI just crossed from "interesting in a lab" to deployable. Chai-1 predicts molecular structure; Chai-2 co-designs sequence and structure together, the jump from reading to writing. Their benchmarks show roughly a 20% average hit rate on binders (molecules that stick to a biological target) across 50 targets, finding at least one binder for about half. Patil adds a useful infrastructure point: NVIDIA's newest chips are tuned for large language models, and protein design models want a very different memory and compute profile, so Chai is competing for scraps on hardware optimized for someone else's problem.

The pattern here travels even if the proteins don't. The defensible asset is proprietary wet-lab data from partners like Eli Lilly, not the architecture, which AlphaFold and Meta's ESM already made public. McPartlon's "one-shot drug-ready candidate" framing is a long way ahead of 20% hit rates.

Full analysis

Your draft

Chai Discovery, a 30-person shop valued at $4B after a $400M Series C, is building AI that designs antibodies from scratch, not just predicts how proteins fold. Matt McPartlon and Neil Patil laid out the whole thesis on Latent Space: structural biology AI just crossed from useful-in-theory to deployable, and the compute market hasn't noticed that these models don't look anything like LLMs.

For a team shipping AI into ad-tech, this is a Type 2 read. Nobody's rewriting their bidder over antibody diffusion models. The value here is the pattern, not the protein: a frontier-vs-open-source argument being tested outside LLMs, a validation loop that's brutally slow, and a compute stack built for the wrong workload. Those three things travel.

The Skeptic. The headline number is a ~20% average hit rate on binders across 50 targets, with binders found for about half. That's genuinely good for antibody design. It is also nowhere near "one-shot drug-ready candidate," which McPartlon floats as where this is heading. A binder is not a drug. Developability, immunogenicity, manufacturability, clinical failure at Phase II for reasons no structure model sees. The 0.33 angstrom RMSD cryo-EM match is one design, cherry-picked by definition. Beautiful demo. The question is the distribution, not the hero shot. For a PM: they can reliably make a molecule that sticks to a target half the time. Whether that molecule becomes medicine is a decade-long open question.

The Researcher. The architecture story is the real content. Chai-1 predicts structure: tokenizer, transformer trunk, diffusion model spitting out 3D atomic coordinates, same family as image diffusion. Chai-2 adds all-atom diffusion to co-design sequence and structure together, which is the jump from reading to writing. That lineage traces straight back to AlphaFold-2 Multimer in 2021 and Meta's ESM protein language models, where Chai CEO Josh cut his teeth. The scaling-law intuition that worked on text worked on protein sequences. For a PM: the same "throw more data and compute at a transformer" recipe that built ChatGPT is now building molecules, and the people doing it learned the recipe at the LLM labs.

The Compute Pragmatist. Patil's claim is the one worth taking to your infra team. NVIDIA's newest silicon, B300s and Vera Rubin, is tuned for LLMs: huge KV caches, 72-GPU NVLink pods, the memory profile of long-context attention. Structural biology models want something else. L³ attention over pair representations, small hidden dimensions, recursive diffusion. Different compute-to-memory ratio entirely. So Chai is fighting hyperscalers and frontier labs for spot-market scraps on hardware optimized for someone else's problem. The lesson for any AI builder: "GPU-rich" is workload-specific. The chip that's cheap for your neighbor's transformer can be the wrong shape for yours, and the market prices the majority workload.

The Open-Source Advocate. Chai is betting the LLM market structure repeats: open-source commoditizes the easy binders, frontier closed models keep most of the value because the whole pie grows faster than the commodity floor rises. Maybe. But protein design has a wrinkle LLMs don't. The architecture is already public, courtesy of ESM and AlphaFold. What open weights can't touch is the partner data: Chai fine-tunes single-tenant variants on Eli Lilly's and Pfizer's proprietary wet-lab results, isolated per customer. So the "frontier captures value" claim is really a "proprietary data captures value" claim wearing a model's coat. For a PM: the defensible asset is the experimental data loop nobody else can see, not the architecture, which ESM and AlphaFold already made public.

The Builder. Patil said something worth writing down about building on top of fast-moving models: software meant to last 20 years now lasts one. Every model generation, Chai-4 and beyond, forces a rewrite of the design suite at a higher level of abstraction, from molecule-inspector to campaign-orchestrator. And they threw out the chatbot UI. It's a CAD tool, "Photoshop for molecules," where computational biologists paint epitopes and content-aware-fill the binder. That's the transferable call. When the model underneath you gains a capability tier every release, you build the thin, rewritable interface, not the cathedral. Anyone shipping LLM features into a product already feels this: the abstraction you hard-code today is technical debt by the next model drop.

The tensions. The Skeptic and the Researcher split on the trajectory: Chai's own benchmarks are real, but McPartlon's "hit-discovery and lead-optimization collapse into one shot" is a bet the data doesn't yet support. The Open-Source Advocate and the Compute Pragmatist agree on where the money is but for different reasons: one says proprietary partner data is the moat, the other says the compute mismatch keeps competitors out because you can't cheaply brute-force your way in on the wrong hardware. And the deepest constraint sits under all of it: McPartlon would remove the wet-lab validation loop "by fiat" if he could. Weeks per validation versus hours for an LLM eval. That gap caps how fast any of these models can self-improve, which is exactly the thing that made LLMs run away.

What it hinges on. For your team, nothing here is a build-or-buy decision this quarter. The one belief worth carrying out: recursive self-improvement needs a fast feedback loop, and every AI domain moves at the speed of its slowest eval. LLMs got hours. Biology gets weeks. That is the whole reason biotech AI won't compound at LLM speed no matter how good the diffusion models get. If you're evaluating any AI capability claim, ad-tech included, ask how long the correctness signal takes to come back. That number sets the ceiling on how fast the system learns.

Prediction: By the time Chai ships its next major model (Chai-4, on the roughly annual Chai-1/2/3 cadence, so before end of 2026), it will still report antibody hit rates in the low tens of percent, not the near-100% "one-shot drug-ready" regime McPartlon gestures at.

Confidence: Medium. The wet-lab validation loop caps improvement speed structurally.

Why: Chai-2 lands around a 20% average hit rate, and each model generation depends on new experimental data to train the next, but that data comes back in weeks, not hours, which McPartlon himself names as the bottleneck he'd remove "by fiat." A capability that self-improves fast needs a fast feedback loop, and biology doesn't have one, so the jump from "binder half the time" to "drug-ready one-shot" is a multi-generation climb, not a single release away. The opposite outcome, near-perfect hit rates in the next cycle, would require the validation loop to stop mattering, and nothing in the conversation suggests that's close.

Revisit by 2026-12-31: We're right if Chai-4 (or Chai's next headline benchmark) reports hit rates in the tens of percent. We're wrong if Chai publishes a validated benchmark showing reliable one-shot drug-ready candidates at 80%+ hit rates.

The binder-is-not-a-drug gap is where the hype and the science part ways. Chai's real product is the partner data flywheel, and that compounds at wet-lab speed, which is slow on purpose.

Comments