Refacto AI

Podcast episode

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

ai-risk biosecurity genomics regulation

Eric Nguyen, CEO of Radical Numerics, joined the Latent Space podcast to make a case that bio-security is now an AI problem. The same models that generate DNA sequences from scratch are the best tools for catching dangerous ones, and right now the builders are ahead of the defenders.

The concrete evidence: a team at Arc Institute and Stanford used Nguyen's Evo model to synthesize a working virus genome, about 6,000 base pairs (DNA letters), in a real lab. His newer model, Omni, tops benchmarks on predicting which human gene variants cause disease. Clément Delangue of Hugging Face argues the offensive capability is already published, already loose, so defensive tooling has to be open too. Nguyen half-agrees. Greg Brockman of OpenAI apparently spent four months of his sabbatical debugging genomic model code at 3am, which tells you how seriously the frontier labs are watching this space.

The working virus is real, but 6,000 base pairs is nowhere near a human genome. The scary demo and the doomsday scenario have a lot of unsolved engineering between them.

Full analysis

Eric Nguyen, CEO of Radical Numerics, went on Latent Space to make a claim that should stop anyone who thinks AI safety is a chatbot problem: the same AI that writes DNA sequences from scratch is also the best tool for catching dangerous ones, and right now the people building detection tools are losing. A team at Arc Institute and Stanford already used the Evo model to generate a working virus genome, about 6,000 base pairs, that got synthesized in a lab and functioned. Nguyen's new model, Omni, is the same family, tuned to answer useful questions and topping benchmarks on predicting which human gene variants cause disease.

For the reader who buys and builds with AI, this is not your stack. But it tells you where the frontier is going, and where the next regulatory wall gets built.

The Skeptic. One virus genome got synthesized and worked. That is real, and it is roughly 6,000 letters of DNA, which is tiny. The human genome is 3 billion. Nguyen's model reads about 2 million at a time. That is a 1,500x gap between what the model handles and a full human genome. The scary demo and the doomsday scenario are separated by a lot of engineering nobody has done yet. On the defense claim, "the same weights detect and generate" sounds elegant, but a detector is only as good as what it was trained against, and Nguyen is describing attackers who deliberately change the spelling to dodge detection. That is a moving target. He is selling the cure and the disease. Note that.

The Researcher. The genuinely new thing here is a model trained only on raw DNA that turned out to be good at protein tasks it was never shown, and now tops human variant-effect benchmarks in the non-coding regions, the parts of the genome that don't make proteins and that older models kept getting wrong. That is a real capability jump, on a real benchmark, in the hard part. The "biological chain-of-thought" claim, showing the model better and better RNA sequences and having it extrapolate to new high-fitness ones, is the interesting part and the unproven part. Wet-lab validation is "in progress." Until those results come back, that is a slide, not a result.

The Open-Source Advocate. This is the whole fight in one story. Clément Delangue of Hugging Face argues that defense capability has to be open to keep pace with attack. Nguyen agrees the design models exist and are demonstrated in public. Evo was on the cover of Science. The bacteriophage result is published. The offense side is already out. So the question is not whether to open-source the dangerous capability, it already leaked into the literature. It is whether the defensive tooling, the detection and attribution, stays locked inside one startup and a national lab, or gets distributed to every DNA synthesis company that needs it. If defense is proprietary and offense is published, defense loses by construction.

The Compute Pragmatist. Watch the money trail, because it tells you who thinks this matters. NVIDIA backed Evo 2. Greg Brockman, sitting president of OpenAI, burned four months of a sabbatical debugging genomic model code at 3am. That is not idle curiosity. Jensen Huang gets a new workload that eats GPUs the way language models do, and biology has 1,500x more context to chew through, so it eats more. The binding constraint here is the same one that gates every long-context model: the architecture work to read millions of DNA letters at once without the cost exploding. That is why the state-space and striped-hyena tricks matter. Whoever cracks full-genome context sets the pace for the whole field.

The Safety Lens. Nguyen threw the punch that lands hardest: he sat on a panel where every AI scientific-discovery company named bioweapons as the top risk, then looked at those same companies and none of them had a real effort in the space. Anthropic and the others flag bio-risk at the chatbot level, they screen the words. Nobody screens the substrate, the actual sequences. And the near-term worry runs to volume, not state actors. Nguyen says the bar for expertise drops and the number of people who can attempt this goes up exponentially. Brandon from Atomic AI named the disanalogy with cybersecurity that should scare people: you can patch software. You cannot patch a human genome.

Where they part ways

The Skeptic and the Researcher split on the same demo. One sees a 6,000-letter proof of concept that is 1,500x short of the nightmare and years of engineering from it. The other sees a model that generalizes across biology in ways nobody predicted, which is exactly the pattern that precedes a fast capability jump. Both are looking at the same evidence.

The Open-Source Advocate and the Safety Lens agree offense is already loose but disagree on the fix. Open the defense to everyone, or concentrate it where it can be governed? If you distribute detection widely, you also hand attackers a map of exactly what gets caught.

And everyone disagrees with the frontier labs by implication. The labs say bio is the number one risk. Nguyen says they don't staff it. That gap is the story.

What it hinges on

Three things decide whether this matters to you in the next year. First, does the "biological chain-of-thought" survive the wet lab? If the model's proposed novel sequences actually work at the bench, that is a genuine design engine and the arms race is real. If they don't, this is a strong prediction tool and a marketing frame. Second, does anyone force DNA synthesis screening to move from pattern-matching to model-based detection? Right now the synthesis companies check new orders against a database of known bad sequences, and Nguyen's whole point is that AI can produce a functionally identical order with different letters that sails through. Third, does the money follow the rhetoric at the frontier labs.

For you, directly: nothing to deploy, nothing to buy this quarter. But if you build anything that touches biology, synthesis APIs, genetic data, hospital genomics, budget for defense-side tooling as a real line item, not an afterthought. And expect a regulatory surface to form here fast, because governments are already engaging Radical Numerics.

Prediction: No US or EU rule will require DNA synthesis providers to use AI-based sequence screening, beyond today's database pattern-matching, before the next US biosecurity policy cycle concludes in March 2027.

Confidence: Medium. The capability is public but the regulatory machinery is slow and the tools are immature.

Why: The offensive capability is already demonstrated and published: Evo generated a working bacteriophage genome, and Microsoft's own "paraphrase" research shows you can keep a sequence's function while changing its letters to dodge detection. That means the vulnerability in current screening is documented, which is usually what precedes a mandate. But mandating AI-based detection requires a detector that regulators trust, and Nguyen himself says the defensive tools are early and one startup plus a national lab are building them. You cannot write a rule requiring a tool that isn't validated and widely available yet. Regulators move at the speed of proven tooling, and the tooling isn't proven. The opposite outcome, a fast mandate, would need both a validated detector and a triggering incident, and neither is on the calendar.

Revisit by 2027-03-29: We're right if DNA synthesis screening rules in the US or EU still rest on database matching with no AI-detection requirement. We're wrong if either jurisdiction issues a rule or binding guidance requiring model-based sequence screening.

Comments