Refacto AI

Podcast episode

πŸ”¬ The Coolest Diffusion Research Isn't in LLMs β€” Evan Feinberg & Sergey Edunov, Genesis Molecular AI

ai-in-adtech cloud-costs measurement

The most transferable idea in this episode has nothing to do with drug discovery: when an AI agent's per-step error stays below a hard physical threshold β€” in PEARL's case, roughly 1 Γ… in protein-ligand pose prediction β€” the loop compounds toward a solution instead of off a cliff. Genesis's founders, both with Meta/Llama pedigree, built that threshold into a diffusion model for 3D molecular structure and got a wet-lab partnership with Incyte to prove it. The "transformers are boring" line is a podcast provocation, not a technical argument β€” but the underlying point is real: inference-time scaling over structural representations, with physics-based verifiers inside the diffusion loop, is fresher territory than the tenth fine-tuning recipe. Watch whether outside labs replicate PEARL's headline numbers; self-reported "beats all public models" from the team selling the model is exactly the evidence to discount.

Full analysis

Two former Llama pretraining leaders now argue the real diffusion frontier is 3D molecular structure prediction, not language β€” and they're backing it with a benchmark-topping drug-discovery model (PEARL), an agentic loop (SAPPHIRE), and a GPU-scarcity warning that spills directly onto everyone else renting H100s.

Reversibility: This isn't a decision for the AI builder β€” it's a signal to read. Type 2 (nothing to commit to). The useful question is what the episode reveals about where architecture innovation, inference-time scaling, and compute pressure are heading.

What's actually being decided: Nothing for ad-tech. The honest framing: does anything here change how a technical AI leader building ad products should think about model architecture, test-time compute, benchmarks, or GPU access? Two of those, yes. Ad-tech relevance specifically: near zero.

Timeline: No forcing function. This is a "where is the frontier moving" read, not a "ship by Friday" call.


The Skeptic β€” The headline claim β€” "LLM architecture has been stagnant since 2017" β€” is the kind of thing that sounds profound on a podcast and dissolves under scrutiny. Flash attention, RoPE, GQA, MoE routing, and the entire reasoning-model regime are not "the same transformer." A co-folding startup has every incentive to say language is boring and molecules are exciting. Second: "correct for every single pose" on 802 complexes of one protease is a narrow, self-selected slice. Zero-shot on a released-after-cutoff template is legitimately good practice, but one target isn't generalization. For an ad-tech reader, the transferable lesson is the only durable one here: your benchmark can pass while production quietly fails.

The Researcher β€” The genuinely instructive piece is the eval critique, not the architecture boast. 2Γ… RMSD (root-mean-square deviation β€” how far predicted atom positions sit from truth) as a pass threshold is exposed as meaningless when hydrogen bonds have a 0.6Γ… tolerance. That's a real methodology point: the community metric was calibrated to what models could hit, not to what the science requires. The move toward lDDT is the correct response. For an informed PM: imagine grading a translation model as "correct" if half the words are right β€” that's the mismatch they're calling out. The inference-time-scaling parallel is also real β€” iterative diffusion refinement with a physics verifier is genuinely the same shape as reasoning tokens with a checker.

The Compute Pragmatist β€” This is the one line that touches every AI builder regardless of domain: GPU scarcity is Genesis's #1 constraint, and they blame LLM demand for crowding them out. NVIDIA invested twice and co-authored their paper β€” meaning Jensen is actively seeding non-LLM demand for the same silicon you're bidding on. Translation for an ad-tech team running real-time bidding models or creative generation: the H100/H200 spot market now has a new class of well-funded, patient, physics-simulation buyers who will happily eat capacity for month-long MD data-generation runs. Your inference bill's ceiling is set by who else wants the chips, and the "who else" list just got longer and richer.

The Open-Source Advocate β€” Note what's not here: PEARL is closed, NVIDIA-entangled, and pharma-partnered. Contrast with the protein world, where ESMFold and AlphaFold weights seeded a whole ecosystem. The episode says ESMFold authors are citing Genesis's architecture papers β€” so the ideas leak even when weights don't. For a builder, that's the pattern to watch: architectural innovations from molecular diffusion (verifier-guided iterative refinement, physics-as-reward) are publishable and portable, even if the specific models stay proprietary. You can borrow the mechanism without the moat.


Sharpest tensions:

  1. Researcher vs. Skeptic on "stagnant LLMs." The Researcher grants that transformer cores are similar lab-to-lab; the Skeptic says the training regime, not the layer, is where the innovation went β€” and calling that "boring" is a founder's rhetorical convenience. Both are right about different layers of the stack, which is exactly why the soundbite is misleading.

  2. Compute Pragmatist vs. everyone. The only claim that reaches into an ad-tech reader's actual P&L is GPU competition β€” and it's the least discussed in the summary. The architecture debate is intellectually fun; the chip-market signal is the one with a dollar figure attached.


What this hinges on: For a digital-advertising / ad-tech audience, the direct impact is low, and it should be stated plainly β€” this is a drug-discovery episode with no ad-tech, publisher, agency, or media-buying implication. The two things worth extracting are portable, not domain-specific: (1) verifier-guided inference-time scaling is now demonstrably a general technique, not an LLM-only trick β€” relevant if you're building anything with iterative refinement and a checkable signal; (2) GPU demand for frontier compute is broadening beyond the hyperscaler labs, which tightens the same market ad-tech ML teams rent from.

Nothing here needs verifying or de-risking for an ad-tech operation. File the eval-quality lesson (benchmark thresholds calibrated to model ability, not to real-world tolerance) as a reusable mental model for your own production evals, and move on.


Prediction: By NVIDIA's GTC 2027 keynote (spring 2027), NVIDIA will publicly spotlight life-sciences / drug-discovery AI (including a Genesis-type co-folding partner) as a named growth segment alongside its LLM and robotics narratives.

Confidence: Medium β€” NVIDIA already invested twice in Genesis and co-authored PEARL.

Why: A chipmaker that co-authors a partner's technical report and invests repeatedly is building a reference customer it will showcase; GTC keynotes are where NVIDIA formalizes new vertical demand stories, and seeding non-LLM GPU demand is directly in its interest.

Revisit by 2027-04-30: We're right if NVIDIA's GTC 2027 keynote or its official GTC materials name drug-discovery/molecular AI as a distinct growth vertical with a named model partner. We're wrong if life-sciences AI gets no dedicated keynote billing and remains folded into a generic "science" mention.

Comments