Refacto AI

Podcast episode

AI:AM Highlights: Astra as AGI, OpenAI's Pause, Mythos @ Mozilla & Human Agency vs Technocapitalism

agents cloud-costs evals inference model-pricing

AI:AM hosts Nathan Labenz and Prakash Narayanan dig into two things this week: whether OpenAI's "pause" on its biggest training runs means anything, and whether GPT-6 Astra qualifies as AGI (artificial general intelligence, meaning a model that can do knowledge work autonomously at human level or above).

The pause, they conclude, is cosmetic. OpenAI kept running inference, kept distilling the big model's behavior into smaller ones, and Astra still took over chunks of OpenAI's own research. On the capability side, the new model solves 25 to 45% of hard math problems versus 10 to 15% for Astra, but burns ten times the compute per answer to do it. Raffi Krikorian ran Anthropic's security tool against Firefox; a full pass would have cost hundreds of thousands of dollars. Mozilla got free credits.

The "pause" costs the labs nothing they weren't already constrained on. The compute bill, though, is moving whether you plan for it or not.

Full analysis

Two frontier labs are now telling the same story: slow down. OpenAI claims it paused the biggest reinforcement-learning runs (training a model by rewarding good answers), and Anthropic's Dario Amodei published a plan to slow development. Hosts Nathan Labenz and Prakash Narayanan pick that story apart and find the pause is mostly cosmetic. That's the thing worth chewing on: what a "pause" is worth when the incentives all point the other way, and what a model that can do a two-week job on its own does to your compute bill.

This is easy to undo on your side. Nobody is asking you to sign anything. What's actually being decided is how much you trust two things: the safety numbers labs publish, and the story that the frontier is slowing. There's no deadline. But if Prakash Narayanan's read on GPT-6 Astra is even half right, your per-user compute costs are about to move whether you plan for it or not.

The Skeptic

A "pause" that stops half of one training method, and even that incompletely, is not a pause. Nathan Labenz and Prakash Narayanan say it plainly: OpenAI kept running inference, kept training smaller models by copying the big one's behavior, and the Astra-class models still took over part of OpenAI's own research work. When a company under an IPO decision publishes a chart that lumps "more capable than Astra" into a reassuring blue bucket, read the chart as marketing. Amodei's slow-down plan and Altman's "let's pace it" land the same week. Both cost these labs nothing they weren't already constrained on, and both make regulators feel handled.

The Researcher

Look at the actual numbers before you buy the AGI headline. OpenAI's next-gen model solves 25 to 45% of curated hard math problems versus 10 to 15% for Astra, but it burns roughly ten times the compute at answer time to get there. That's a real jump and an expensive one. The more interesting line comes from Nathan Labenz: the Meter benchmark, which measures how long a task an agent can finish on its own, is going obsolete because models now improve faster than the tasks take to run. When your yardstick can't keep up with the thing it measures, published capability numbers stop meaning what buyers think they mean.

The Safety Lens

Apollo Research got three days to evaluate Astra before release. That's the whole governance story in one number. The auditors checking these models are funded by, and hire from, the labs they audit. Redwood Research pays technical staff $350K to $850K; the talent flows one direction, toward the labs. Paul Christiano, now on OpenAI's safety committee, says out loud he expects we "permanently lose control" of superintelligence built without better alignment. For anyone buying AI on the strength of a published safety evaluation for compliance or procurement: treat those documents as thin. A three-day look is not due diligence you can lean on.

The Compute Pragmatist

Prakash Narayanan's line is the one that hits your budget: OpenAI can afford to pace because it has the compute; Anthropic can't, so it needs better models to compete. That means the pace of the frontier gets set by whoever is in second and third place and hungry, not by the leader's restraint. And when Astra can autonomously finish a 1.5-to-3-week job one time in six, and two-thirds of the time with a human nudging it, the compute per task goes up hard. Raffi Krikorian at Mozilla ran Anthropic's security tool against Firefox and a full run would have cost "hundreds of thousands of dollars." Mozilla got free credits. You won't.

The Builder

The capability is real enough to build on, and messy enough to hurt you. Astra's "computer use" (the model operating a browser or desktop itself) now works reliably inside a set time window, and a persistent notes file lets it carry context past the million-token limit. Good. But the code it writes is inconsistent: high quality on GPU kernels, sometimes functional-but-unreadable elsewhere. Raffi Krikorian says Mozilla's Firefox team still requires a human to review every commit; their AI team lets agents rewrite the codebase hourly as long as tests pass. That split is your decision too. Where a bad output is expensive to unwind, keep the human on the commit. Where tests catch the failure cheaply, let it run.

Where they part ways

The Compute Pragmatist and the Skeptic disagree on what the pause even is. The Pragmatist reads it as a real constraint on the have-nots, Anthropic pacing because it can't afford not to. The Skeptic reads both labs' slow-down talk as theater timed to the regulatory and IPO moment. They can both be right: Anthropic is genuinely compute-starved, and the public messaging is still engineered to reassure.

The Researcher and the Builder split on whether "AGI" means anything you can act on. The Researcher sees a model that pays ten times the compute for a math score and warns the benchmarks are breaking. The Builder shrugs and ships, because a model that finishes real multi-week work two-thirds of the time is worth paying for even if the label is hype.

What this actually hinges on

One belief: is the frontier slowing, or just the messaging? Everything downstream, your compute budget, how much you trust safety evals, whether you staff for human-in-the-loop, turns on that. The council leans hard toward "the messaging slowed, the work didn't." Christiano's own words, the three-day audit, the incomplete pause, and Narayanan's point that the real frontier lives in researchers' heads twelve to eighteen months ahead of any shipped model all point one way.

What to verify before you lean on any of this: run your own eval on the specific task you care about, because the public benchmarks are going stale. And if you buy AI where a safety or security claim matters, ask the vendor how many days the external auditor actually got. If the answer is three, you know what it's worth.

Prediction: Before OpenAI's next frontier model release after GPT-6 Astra, no US or China frontier lab (OpenAI, Anthropic, Google DeepMind, Meta, xAI) will agree to a binding, externally verifiable cap on frontier training compute, and each will keep shipping more capable models on its normal cadence.

Confidence: High. The incentive to keep training points the exact opposite way from the pause talk.

Why: OpenAI's own "pause" stopped only part of one training method and was incomplete even then, while its next-gen model already beats Astra on hard math. Anthropic is compute-starved and, as Prakash Narayanan argues, the hungry runner-up sets the pace, with the leader's stated restraint irrelevant to that dynamic. A binding cap would require Sam Altman, Dario Amodei, Mark Zuckerberg and Elon Musk to all agree while each fears the others defecting, and the podcast notes even military hotlines between the US and China go unanswered. The opposite outcome, a real verifiable cap, needs trust and enforcement machinery that does not exist and that no lab has any reason to build first.

Revisit by 2027-03-18: We're right if every named lab has shipped at least one more capable model and none has signed a binding, third-party-verified training-compute cap. We're wrong if any of the five agrees to an externally audited limit on frontier training compute.

Comments