Refacto AI

Industry story

AI Safety Conference Reveals Labs Have No Alignment Plan

agents alignment evals safety

At The Curve conference (its third annual iteration), a gathering of AI safety researchers, lab employees, policymakers, and investors, the dominant takeaway was that no frontier AI lab has a credible plan for aligning systems smarter than current models. The described 'plan' amounts to: solve near-term alignment problems with operational excellence, then use AI to do the harder alignment work automatically — a process the author and others call 'No Plan,' noting it relies on automated recursive self-improvement (RSI, where AI iteratively improves itself) that experts consider extremely dangerous. Labs are reportedly aware this is insufficient and are worried they may not even have time to execute this inadequate approach, given how fast capabilities are advancing.

Analysis

Showing the shorter version.

The Zvi's writeup of The Curve conference lands on one finding: no frontier lab has a credible plan for controlling AI smarter than today's models. The stated plan, named plainly by people inside OpenAI and Anthropic, is to solve the easy alignment problems by hand and hand the hard ones to a future AI that improves itself. The biggest question senior staff faced was "why are you doing this," and the answers left the room unsatisfied.

The Skeptic's objection is fair as far as it goes. A safety conference concluding there's no safety plan is not a surprise. "No Plan" also flattens a real distinction: inadequate is not the same as nonexistent. But the Safety Lens gets the more important point. The labs that have built their commercial brands on safety leadership are the ones saying this at a microphone. That's different from researchers complaining from the outside. The decision to keep shipping rests on competitive pressure, not on a safety case anyone in that room could defend out loud.

The scheduling problem is concrete. Capability jumps track training-run cadence, measured in months. Interpretability, formal verification, and scalable oversight, the three lines of work that would actually matter, are measured in years and are already behind where they'd need to be for models arriving next year. More compute accelerates the wrong half of that gap. You cannot buy insight by the GPU-hour.

For buyers, this doesn't show up in a procurement checklist. You sign for SOC 2, data residency, and indemnification. There's no line item for "has a plan to control a model smarter than the one you're licensing." The labs named the gap and will keep shipping into it, because their deadline is set by each other. So the practical move is narrow: discount safety-as-marketing by a notch, run your own evals, and when a vendor's pitch leans on safety, ask what specifically binds that promise beyond a blog post. Don't treat a vendor's safety posture as a substitute for your own testing.

Prediction: By The Curve's next annual conference in autumn 2027, neither OpenAI nor Anthropic will have published a technical alignment plan for superhuman systems that does not rest on AI-assisted alignment. Confidence: medium. The incentive to keep shipping is stronger than the incentive to publish a plan they'd then miss publicly. We're wrong if either lab publishes a superhuman-alignment plan whose central mechanism is human-driven interpretability or scalable oversight that does not route through automated self-improvement. Revisit by 2027-11-30.

Also covered this issue

Comments