Podcast episode
Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week
agents inference model-pricing open-weights tool-use
Ari Weinstein and Nikunj Handa joined the podcast to walk through what OpenAI actually shipped at DevDay: a cheaper model (GPT-4.1 Sol), a personal assistant called Dots with its own cloud computer, a Computer Use tool that controls a browser the way a human would, and a Decisions API that classifies requests fast. That last one is a near-copy of a startup called Jev, built in about a week.
The substance that matters: Computer Use got 7x faster because OpenAI rewrote it to fire multiple browser actions per step instead of one screenshot at a time. The Decisions API is fast but ships with no calibrated confidence score, meaning you can't reliably threshold it. Handa admitted training pushes it toward answers people like, not answers that are right. For triage, probably fine. For anything that touches money, that's a real problem.
The infrastructure pieces, caching, async calls, history compaction, are worth adopting now. The autonomous agent loop is not ready for production work where errors are expensive.
Full analysis
OpenAI used DevDay to turn its internal agent plumbing into things you can buy. A cheaper model (GPT-6.1 Sol), a personal assistant with its own cloud computer (Dots), a Computer Use tool in the Agents API, and a fast-classification endpoint (the Decisions API) that OpenAI copied from a startup called Jev in about a week. The through-line: the tricks that made OpenAI's own products fast and cheap are now rented out to everyone else.
What's actually being decided, for someone running an AI product: do you keep building your own agent scaffolding and your own fast-decision layer, or do you move onto OpenAI's native stack because it's cheaper and better-aligned to their models? That's easy to undo at the margins and hard to undo if you rebuild your whole agent loop around it. No deadline here. This is a repricing event, and repricing keeps coming.
The Skeptic
"Faster than the average human" is a benchmark claim, and the average human is a low bar for clicking through a web form. Weinstein wants expert-speed next, which is the hard part and the part he's not showing numbers for. Watch the gap between the demo and your actual traffic: Computer Use still fails on the long tail of weird UI states, and the saved paper on agents in live, changing environments exists precisely because static benchmarks lie. The Decisions API is "zero-shot" with no calibrated confidence. Nikunj Handa admitted RLHF (training a model to give answers people like) pushes it toward what you want to hear, not toward being right. So you get fast decisions with no trustworthy certainty score. For triage, fine. For anything that moves money, not yet.
The Researcher
Almost none of this is a new model. The Decisions API runs on existing Luna weights. No training run. Handa was blunt: "this is really just Luna" plus inference-stack tuning, structured outputs, and parallel batching. The ~80% Luna price cut came from efficiency work before the speed push. That matters because it tells you where the gains are coming from now: not bigger brains, but squeezing the serving layer. Computer Use did change architecture, from one-screenshot-one-action to writing JavaScript that fires many actions per step, plus accessibility trees and DOM access. That's real and it's why the speed jumped 7x. But the capability frontier moved less than the engineering did.
The Open-Source Advocate
Here's the uncomfortable part for anyone running Llama, Qwen, or Mistral behind their own agent. OpenAI's pitch is distribution lock-in, said plainly by Weinstein: they train the model on their own Computer Use harness, so their harness is faster, cheaper, and more accurate on their model than anything you build. If that holds, "bring your own open model plus your own scaffolding" carries a structural penalty you can't engineer away, because you don't control the training. The counter: a Jev clone built in a week on existing weights means the decisions-layer moat is thin. Any lab, open or closed, can ship fast structured classification with inference tricks. The open ecosystem can match the Decisions API. It cannot easily match a harness co-trained with the model.
The Compute Pragmatist
Follow the silicon. OpenAI's prior fast mode, 5.3 Spark, was publicly on Cerebras chips (a non-NVIDIA inference specialist). For UltraFast, OpenAI "declined to confirm or deny" Cerebras. That dodge is the interesting bit. If the speed wins are riding on specialized non-NVIDIA hardware, then the cost curve you're being sold depends on a supply chain OpenAI controls and you don't. The 12-hour cache window, pre-warming, and compaction (automatically trimming the conversation history fed back to the model) are all ways to stop paying to reprocess the same tokens. Real savings for long-running agents. But they tie your cost structure to OpenAI's infrastructure decisions. One pricing change and your unit economics move.
The tensions
The Open-Source Advocate and the Researcher collide on the moat. The decisions layer is cheap to clone, so no moat there. The co-trained Computer Use harness is a real moat, because you can't replicate training you don't own. Both are true, and they point opposite directions for a buyer.
The Skeptic and the Builder-minded reader split on timing. The infrastructure (caching, async tool calls, compaction) is safe to adopt now and saves money. The autonomous Computer Use agent is not safe for money-touching work until the calibration problem is solved, and Handa flagged that as unsolved.
What it hinges on
Does the distribution-alignment advantage actually show up in your numbers? That's the one belief everything rests on. If OpenAI's native Computer Use is meaningfully more accurate on their model than your harness, the lock-in is real and you should plan around it. If it's a few points, keep your options open. Run your own head-to-head: your hardest 50 workflows, OpenAI's Agents API Computer Use versus your current setup, measured on task completion and cost per completed task, not on their benchmark. Take the cheap infrastructure wins now. Hold the line on autonomous decisions until there's a confidence score you can trust.
Prediction: At OpenAI's next major model release after DevDay 2026, the Decisions API will still ship without a calibrated, documented confidence score that developers can threshold on in production.
Confidence: Medium. OpenAI named calibration as unsolved future work, and the fix requires training, not inference tuning.
Why: Handa said the Decisions API is zero-shot on existing Luna weights with no new training, and that RLHF collapses a model toward agreeable answers rather than honest probabilities. Calibrated confidence is a training problem, and OpenAI explicitly put it in the "future work" bucket while shipping the product in a week on inference tricks. Labs ship the fast, cloneable part first and leave the hard training work for a later cycle, because the demo closes deals and calibration doesn't demo well. The opposite outcome, a trustworthy confidence score arriving in the next release, would require OpenAI to prioritize a quiet reliability feature over the next capability headline, which is not how DevDay-driven roadmaps run.
Revisit by 2027-06-30: We're right if the next GPT model release ships a Decisions API whose docs still lack a calibrated confidence output developers can set a threshold on. We're wrong if OpenAI documents and exposes a calibrated confidence score for the Decisions API by then.
Comments