Industry story
AI Safety Conference Reveals Labs Have No Alignment Plan
At The Curve conference (its third annual iteration), a gathering of AI safety researchers, lab employees, policymakers, and investors, the dominant takeaway was that no frontier AI lab has a credible plan for aligning systems smarter than current models. The described 'plan' amounts to: solve near-term alignment problems with operational excellence, then use AI to do the harder alignment work automatically — a process the author and others call 'No Plan,' noting it relies on automated recursive self-improvement (RSI, where AI iteratively improves itself) that experts consider extremely dangerous. Labs are reportedly aware this is insufficient and are worried they may not even have time to execute this inadequate approach, given how fast capabilities are advancing.
Analysis
Showing the shorter version.
The Zvi's writeup of The Curve conference lands on one finding: no frontier lab has a credible plan for controlling AI smarter than today's models. The stated plan, named plainly by people inside OpenAI and Anthropic, is to solve the easy alignment problems by hand and hand the hard ones to a future AI that improves itself. The biggest question senior staff faced was "why are you doing this," and the answers left the room unsatisfied.
The Skeptic's objection is fair as far as it goes. A safety conference concluding there's no safety plan is not a surprise. "No Plan" also flattens a real distinction: inadequate is not the same as nonexistent. But the Safety Lens gets the more important point. The labs that have built their commercial brands on safety leadership are the ones saying this at a microphone. That's different from researchers complaining from the outside. The decision to keep shipping rests on competitive pressure, not on a safety case anyone in that room could defend out loud.
The scheduling problem is concrete. Capability jumps track training-run cadence, measured in months. Interpretability, formal verification, and scalable oversight, the three lines of work that would actually matter, are measured in years and are already behind where they'd need to be for models arriving next year. More compute accelerates the wrong half of that gap. You cannot buy insight by the GPU-hour.
For buyers, this doesn't show up in a procurement checklist. You sign for SOC 2, data residency, and indemnification. There's no line item for "has a plan to control a model smarter than the one you're licensing." The labs named the gap and will keep shipping into it, because their deadline is set by each other. So the practical move is narrow: discount safety-as-marketing by a notch, run your own evals, and when a vendor's pitch leans on safety, ask what specifically binds that promise beyond a blog post. Don't treat a vendor's safety posture as a substitute for your own testing.
Prediction: By The Curve's next annual conference in autumn 2027, neither OpenAI nor Anthropic will have published a technical alignment plan for superhuman systems that does not rest on AI-assisted alignment. Confidence: medium. The incentive to keep shipping is stronger than the incentive to publish a plan they'd then miss publicly. We're wrong if either lab publishes a superhuman-alignment plan whose central mechanism is human-driven interpretability or scalable oversight that does not route through automated self-improvement. Revisit by 2027-11-30.
A conference of AI safety researchers concluded that no frontier lab has a credible plan for controlling AI smarter than today's models. The Zvi, writing up The Curve conference, calls the actual plan "No Plan": solve the easy alignment problems by hand, then hand the hard ones to a future AI that improves itself. For a reader who buys and builds on OpenAI and Anthropic, the question is not whether the sky is falling. It's what this admission tells you about the vendors you depend on, and whether anything you do this quarter should change.
This is easy to undo on your side. Nobody is asking you to rip out Claude or GPT. What's actually being decided is how much weight to put on the labs' own safety marketing when you pick a vendor, write a contract, or plan a deployment on longer-running agents. There's no deadline here. No shutdown date, no price change. So treat this as a slow signal, not a fire alarm.
The Skeptic. A room full of safety researchers deciding there's no safety plan is the most predictable result in the world. That's their whole job. "No Plan" flattens a real distinction: an inadequate plan is not the same as no plan, and "inadequate" is carrying the entire argument. The implied counterfactual, that a sufficient plan exists and the labs are ignoring it, is nowhere in the write-up. Nobody at The Curve stood up with the better plan. Alarm about recursive self-improvement is fair. The suggestion that a clear fix is sitting on a shelf being neglected is not. Vivid conference quotes crowd out the dull baseline: alignment has never been solved, and today changes nothing about that.
The Safety Lens. What's new is who said it. The admission came from people inside OpenAI and Anthropic, the two labs that market safety hardest. The biggest question senior staff got was "why are you doing this," and the answers left the room unsatisfied. That means the decision to keep going rests on competition, not on a safety case anyone could defend out loud. Recursive self-improvement as the fallback is the exact scenario every serious risk analysis ranks worst: you're trusting a smarter system to share your goals before you've shown it does. For policymakers, the lesson is blunt. The voluntary-commitment framework has no teeth when the labs writing the commitments are the ones admitting the plan is thin.
The Researcher. The field has known this for years. Watching it named plainly at a major conference is the change. "Use AI to solve alignment" is not a research program. It's an IOU backed by nothing. The three lines of work that would actually matter at the frontier, interpretability (reading what a model is doing inside), formal verification, and scalable oversight, are all years behind where they'd need to be for models arriving next year. The lab roadmap, stated plainly, is: we will get lucky, or a sub-field we don't yet have will crack it in time. Hope is not a timeline.
The Compute Pragmatist. The admission that labs may not have time to run even the inadequate plan is a scheduling fact, not a philosophy. Capability jumps track training-run cadence, measured in months. Alignment research runs on human insight and interpretability tooling, measured in years. More compute speeds up the capability side far more than the alignment side, because you can't buy insight by the GPU-hour. Every new frontier cluster widens the exact gap The Curve just named. The constraint was never ideas. It's serial iteration speed, and money accelerates the wrong half.
The Enterprise Buyer. None of this shows up in a procurement checklist, and that's the problem. You sign for SOC 2, data residency, SSO, and indemnification. There is no line item for "has a plan to control a model smarter than the one you're licensing." The useful move is narrow: when a vendor's deck leans on safety as a selling point, ask what specifically binds that promise beyond a blog post. Ask for the audit log, the eval results on your own data, the rollback story. The abstract alignment question won't help you close a deal. The concrete "show me, don't tell me" version protects you from paying for marketing.
The council splits on two things. The Skeptic and the Safety Lens disagree on what the admission is worth: the Skeptic says it's the sound a safety conference always makes, the Safety Lens says the source makes it different, because the labs selling safety are the ones confessing. Both are right about different halves. It's not news that alignment is unsolved. It is news that the people who built their brand on solving it will say so at a microphone. The second split is Researcher versus Compute: is the bottleneck ideas or time? They converge on the uncomfortable answer. It's both, and compute spending makes the time problem worse while doing nothing for the idea problem.
What this hinges on for a buyer is simple. Does any of this change a vendor's behavior, or just their conference rhetoric? The incentive answer is that it changes nothing operationally. The labs named the gap and will keep shipping into it, because the thing setting their deadline is each other, not safety. So the practical take is to discount safety-as-marketing by a notch, run your own evals, and don't sign anything that treats a vendor's safety posture as a substitute for your own testing.
Prediction: By The Curve's next annual conference in autumn 2027, neither OpenAI nor Anthropic will have published a technical alignment plan for superhuman systems that does not rest on automated AI-assisted alignment (the "use AI to solve it" approach named at the 2026 conference).
Confidence: Medium. The incentive to keep shipping is stronger than the incentive to publish a real plan.
Why: The Curve's 2026 takeaway, per The Zvi, is that the labs' own staff described the plan as solving easy problems by hand and handing the hard ones to a future AI, with the "why are you doing this" question going unanswered. Nothing in the story suggests a substitute plan exists anywhere in the field, and the Researcher's read is that interpretability and scalable oversight are years behind the capability curve. The deadline for each lab is set by its competitors, so additional compute buys faster capability rather than faster alignment insight, and publishing a hard plan would mean committing to a standard they'd then miss publicly. The opposite outcome, a genuinely new published plan, would require a research breakthrough nobody at The Curve claimed to have.
Revisit by 2027-11-30: We're right if, as of The Curve's 2027 conference, the published alignment approach from both OpenAI and Anthropic still depends on AI doing the core alignment work. We're wrong if either lab publishes a superhuman-alignment plan whose central mechanism is human-driven interpretability, formal verification, or scalable oversight that does not route through automated self-improvement.
Also covered this issue
-
Jay Clayton Named US AI Czar; Administration Posture Shifting
zvi-vase
A new AI regulator may soon treat model providers like chipmakers, reshaping who you can legally buy inference from overseas and what compliance costs land on your vendor contracts.
-
AI Milestone: Fields Medal-Level Discovery and Navier-Stokes Solved in 2026
zvi-vase
Unverified claims about AI solving century-old math problems should not change your infrastructure spending until a proof appears that independent mathematicians can actually check.
Comments