Refacto AI

Podcast episode

Zero to One in AI Safety: Halcyon's Mike McCormick on Launching 30 New Orgs & the Founder Bottleneck

ai-safety interpretability open-weights security

Halcyon co-founder Mike McCormick joins Erik Torenberg and Nathan Labenz on Torenberg's podcast to explain what Halcyon does: it seeds new organizations working on AI safety, biosecurity, and cybersecurity, about 30 so far in three years, partly as a grant-maker and partly as a venture fund.

Two things McCormick describes could land in your vendor stack within a few years. The first is AIUC, a safety certification paired with an insurance policy, modeled roughly on SOC2 (the security audit enterprise procurement already demands). ElevenLabs used it to unlock Fortune 1000 accounts. The second is interpretability as a paid vendor category: Goodfire raised over $100M at a $1B+ valuation to commercialize the practice of reading a model's internal wiring to understand why it produced a given output.

McCormick is honest about the hole, though. Open-weight models, ones whose parameters anyone can download and modify, have no safety solution he'd fund. That admission is more useful than the 30-org headline. The certification market is real, but the enforcement layer underneath it is, by his own description, totally nascent.

Full analysis

Mike McCormick runs Halcyon, a mix of charity grant-maker and venture fund that has helped start about 30 AI safety, biosecurity, and cybersecurity organizations in three years. His pitch: the thing holding AI safety back is not money, it's a shortage of people who can build companies. For the reader who buys and ships AI, the practical takeaway is buried under the philosophy. Two things in this episode could show up in your procurement and vendor list within a couple of years: AI safety certification you'll be asked to hold, and interpretability audits becoming a paid vendor category.

This is easy to react to and hard to get wrong by ignoring. Nothing here forces a decision this quarter. Both commercial patterns will change what you buy and what your customers demand of you.

The Skeptic

Most of this episode is a fundraiser's worldview, not news you can act on. "The bottleneck is founders, not capital" is exactly what a person sitting on a pile of capital says. It flatters his donors and explains why the money isn't producing results yet. Notice the timeline hedge too: McCormick prioritizes work "relevant in the next 1, 2, 3, 4 years," but if takeoff is 3-6 months away "all bets are off." That's a claim that can never be wrong. On the concrete stuff, he's honest about the hole: open-weight model security, where anyone can download and modify the parameters, has no working solution he'd fund. That admission matters more than the 30-org headline.

The Enterprise Buyer

AIUC is the one thing here a CTO could sign. It pairs a safety standard (think SOC2, the security audit enterprise buyers already demand from vendors) with an insurance policy you can only buy if you pass the standard. ElevenLabs used the certification to win Fortune 1000 accounts. That's the model to understand: safety certification becomes a sales unlock, not a compliance cost. If you sell AI features into big companies, expect their procurement teams to start asking whether you hold something like AIUC1. And if you're the buyer, a vendor carrying real insurance against AI failure is a cleaner signal than a page of safety promises.

The Researcher

Goodfire raising $100M+ at a $1B+ valuation tells you that interpretability, which means reading a model's internal wiring to see why it produced an output, just got priced as a standalone business, not a lab research cost center. McCormick calls Goodfire the best place in the world to do this work, Anthropic aside. Take that with a grain of salt from an investor, but the money is real. The other genuinely new item is GRAM from AE Studio: a way to strip dangerous capabilities out of open-weight models by isolating them in specific "experts," the sub-networks inside a mixture-of-experts design where different parts handle different tasks. If it holds up, Meta could ship an open model that structurally can't do the worst things, instead of bolting on guardrails anyone can remove.

The Compute Pragmatist

The verification thread is where the money and the physics meet. McCormick wants zero-knowledge proofs, a cryptographic trick that lets you prove a computation ran a certain way without revealing the data, applied to AI inference, plus "verifiable neoclouds" that can prove what they ran. His own words: the field is "totally nascent." That means the overhead is unsolved. Proving what happened inside an inference run, cheaply, at production volume, is a hard cost problem nobody has closed. Until someone shows the overhead is tolerable, every pacing agreement between labs or nations he's excited about is unenforceable. He knows it. That's why he's convening 25 founders instead of pointing at a product.

The Open-Source Advocate

Here's the uncomfortable part for anyone building on Llama, Mistral, or Qwen. McCormick can't find a fundable answer to open-weight safety, and he'd love to. Once weights are public, guardrails come off. GRAM is the first credible move in the other direction, and Transluce's report on OpenAI agent swarms crawling private databases, covered by TechCrunch this week, shows the threat isn't theoretical. But GRAM only helps if the model maker applies it before release. The reader running open models today carries a risk that neither research nor any vendor has closed. That's not a reason to avoid open weights. It's a reason to know that "someone will fix the safety layer" is not currently true.

Where the tensions are

The Enterprise Buyer sees a product to sign in AIUC; the Skeptic sees a young insurance line pricing a risk nobody can measure yet, which is a fair worry about whether the coverage means anything when a real incident hits. The Researcher sees interpretability priced as an industry at $1B; the Compute Pragmatist sees the enforcement layer underneath it all, verification, admitted to be nascent, which means the certificates and audits rest on trust, not proof. And the Open-Source Advocate and Researcher agree GRAM is the most interesting technical item, while both know it dies if Meta and the other open-model shippers don't bother to use it.

What it actually hinges on

Two questions decide whether any of this touches your business. Does AI safety certification become something your enterprise customers demand, the way SOC2 became table stakes? And does interpretability graduate from research into a vendor category you procure, like penetration testing? AIUC with a real reference customer and Goodfire at unicorn scale are the early evidence for both. Neither is settled. If you sell AI to large enterprises, the cheap move now is to watch which safety standard your biggest customers start naming in RFPs, because that's the one you'll end up holding.

Prediction: At least one Fortune 500 enterprise will publicly name AIUC certification (AIUC1) as a procurement requirement or preferred qualification for AI vendors by 2027-06-30.

Confidence: Medium. One reference customer exists; the pull is real but early.

Why: AIUC already has ElevenLabs using its certification as a sales unlock into large accounts, which means the standard has crossed from idea to a working commercial motion. Enterprise procurement follows a copy-the-leader pattern: once one AI vendor wins deals by carrying a safety credential, buyers start writing that credential into their vendor requirements, exactly how SOC2 spread from a nice-to-have to a checkbox. The opposite outcome, where AIUC stays a single-vendor curiosity, is the less likely path only if a rival standard from a bigger name (a cloud provider or an established audit firm) crowds it out first, which is possible but slower to materialize than AIUC's existing head start. The weak point is timing: enterprise RFP cycles are long, and "by mid-2027" may be early for public naming even if the direction holds.

Revisit by 2027-06-30: We're right if a Fortune 500 company publicly names AIUC1 in a vendor requirement, RFP, or procurement statement. We're wrong if no Fortune 500 buyer has publicly referenced AIUC certification by that date.

Comments