Podcast episode
Why GPT-6 Astra Is So Significant and So Confounding
agents evals model-pricing security tool-use
Nathaniel Whittemore's show takes on GPT-6 Astra, OpenAI's new "computer use" model, meaning software that operates your machine autonomously: it clicks, types, opens apps, and runs for hours without you touching the keyboard. OpenAI is calling it the world's best at this, and the episode tries to figure out whether that's true and whether it matters.
The substance is in the tension between two numbers. On AutomationBench, Astra hits 41.1% versus Anthropic's competing model at 31.4%. That's a real lead. It's also a 59% failure rate on automated desktop tasks, which means you're supervising constantly. The security picture is wilder: Astra scores 100% on ExploitBench, finding and exploiting software holes at a rate seven times its predecessor. That's why OpenAI launched through a cybersecurity partner program rather than dropping it wide open. Guest Thibault adds that the team pulled six months of roadmap forward for Dev Day.
The cost math is the part most people will skip past and shouldn't. Hours of autonomous sessions mean thousands of API round trips, and you pay for the failed ones too.
Full analysis
OpenAI shipped GPT-6 Astra over a weekend and is calling it the world's best "computer use" model. Computer use means the model drives your machine itself: it clicks, types, opens apps, runs a browser, and works for hours without you touching the keyboard. That is the whole story here. Astra is not much better at writing or coding than what you already have. It is better at doing things for you. That gap is why the reviews came in confused, and it is what an operator needs to understand before deciding whether any of this changes the tools they buy.
What's being decided: whether "the model operates the computer for you" has crossed from party trick to something you'd put in a real workflow. This is easy to undo. Nobody is signing a contract. You add an API call or you don't. The deadline pressure, if any, is OpenAI's Dev Day, where their product lead Thibault says they pulled six months of roadmap forward.
The Skeptic One benchmark is doing all the persuading. AutomationBench: Astra 41.1%, Anthropic's Fable 5.1 at 31.4%. Sounds like a rout. It also means Astra fails the automated computer task 59% of the time. A model that gets six of ten desktop workflows wrong is not "hands off my computer." It is a model you supervise constantly, which is the opposite of the launch video's pitch. And the timing is loud. Astra scored 61 on the independent Artificial Analysis index, tied with the old GPT-5.6 Sol and behind Meta's Muse Spark. Then Artificial Analysis rewrote the index over the same weekend and Astra jumped to second. The benchmark moved to fit the model. Read that carefully.
The Researcher Two numbers here are real and one is theater. Real: ExploitBench 100%, and 39% on recently-disclosed security flaws versus 5.5% for the prior model. That is a genuine capability jump in finding and exploiting software holes, and it is why OpenAI gated the launch through a cybersecurity partner program first. Also real, and quietly damning: OpenAI says performance on several tests peaked at "high" effort and dropped at maximum effort. Make the model think harder and it does worse. That means it over-reasons. The theater is Greg Brockman calling this "real AGI." Marcus-on-AI ran two warnings the same day, including Terence Tao's, about exactly this pattern of claim outrunning evidence. Treat the AGI line as marketing.
The Open-Source Advocate Notice what is missing from this conversation. Every comparison is OpenAI's Astra against Anthropic's Fable. No open-weights model within shouting distance of AutomationBench. Computer use is the one capability where the open ecosystem is furthest behind, because it needs long, expensive training on real interface actions, not scraped text. Meta shipped Muse Spark the same weekend and it is a consumer agent, not open weights you can run yourself. So the practical read: if you want an AI that operates a machine, you are renting it from a frontier lab for the foreseeable future. No Llama-shaped escape hatch. That is a dependency worth pricing now.
The Compute Pragmatist Hours-long autonomous sessions are a different cost animal than a chat reply. A prompt is one round trip. Astra driving a CRM for three hours is thousands of round trips, each one a screenshot in and an action out. That is the token bill of a long agent loop, and it runs whether the task succeeds or not. At 41% success, you pay full freight for the 59% that fail too. And the "declines at maximum effort" finding matters to your wallet: cranking the compute setting up can cost more AND deliver worse results. The economics only work where the task is repetitive, high-value, and cheap to verify. Not real-time anything.
The Builder What would I actually ship Monday? Nothing customer-facing that touches money. The security number kills that. A model at 100% on ExploitBench, given a browser and your credentials, is a liability if a prompt gets hijacked. The Claude-token-theft story TechCrunch ran the same day, where attackers drained accounts through compromised integrations, is the preview. Where I would use Astra: internal, sandboxed, human-in-the-loop batch jobs. Data entry, pulling reports across five old web apps, the boring stuff nobody built an API for. Reviewers say front-end design is still Fable's, and Astra overcomplicates simple interfaces. So keep two models. Astra for doing, Fable for making.
Where they part ways:
The Skeptic and the Researcher agree the number is real but read it opposite ways. Researcher: 41% on a brand-new, hard task is a real lead. Skeptic: 41% means you can't trust it unsupervised, so the "ambient computer use" future in the video is years off. Both are right, and the gap between them is exactly the gap between a demo and a deployment.
The Builder and the Compute Pragmatist see the same wall from two sides. The Builder won't ship it near revenue because of the security surface. The Pragmatist won't ship it near real-time because of the cost of long loops. Add those up and Astra's real home is narrow: internal, batch, supervised, high-value-per-task.
And the loudest signal nobody in the launch wants to dwell on: the independent index had to be rewritten in 48 hours for Astra to look like a winner. Either the old benchmarks were genuinely blind to computer use, or the scoreboard bent to the launch. For anyone using Artificial Analysis or similar to pick models, that is the thing to check. Your eval suite may not measure what you actually need automated.
What it hinges on: does 41% computer-use accuracy climb fast, or does it stall like self-driving did at "90% is easy, the last 10% is a decade"? Everything else follows from that one line.
Where the council leans: real capability, oversold readiness. The security jump is the most consequential fact here, and it cuts against enterprise adoption, not for it.
What to verify before you build: run your own five workflows through Astra's computer use and count actual completions without human touch. Run your benchmark, not the vendor's. If you get above 70% on your tasks, you have something. Below that, it is a supervised assistant, priced like an autonomous one.
Prediction: By OpenAI's next flagship launch after GPT-6 Astra (the generation following Astra, expected within the coming year), OpenAI will publicly restrict, gate, or add mandatory guardrails to Astra's autonomous computer-use or exploit capabilities in response to a real-world security incident or abuse report, rather than loosening them.
Confidence: Medium. The 100% ExploitBench score plus a live token-theft pattern makes misuse near-certain.
Why: Astra scored 100% on ExploitBench and 39% on recently-disclosed security flaws, seven times the prior model, and OpenAI itself staged the launch through a cybersecurity-focused partner program before going broad, which tells you they already know the abuse risk is live. The same week, TechCrunch documented attackers draining Claude accounts through compromised integrations, so the attack pattern against agentic models is not hypothetical. A model that can operate a computer and write working exploits, handed to every paid subscriber, will be pointed at something it shouldn't be within months, and OpenAI's own safety-review framing gives them the script to clamp down publicly. The opposite outcome, them widening access untouched, would require zero notable incidents from a capability they themselves flagged as needing extra review, which is the less likely world.
Revisit by 2027-03-14: We're right if OpenAI announces new restrictions, gating, refusal behavior, or usage limits on Astra's computer-use or security capabilities, or ties them to a disclosed incident. We're wrong if access broadens with no new guardrails and no incident-driven walk-back.
Comments