Industry story
OpenAI launches Decisions API to compete with TypeSafe's Jev classifier
agents cost-compression guardrails inference model-pricing
At OpenAI's Dev Day, CEO Sam Altman announced a new 'Decisions API' that gives OpenAI's Luna model a predefined set of options to choose between — such as image categories or agent behaviors — optimized for speed and low cost. The product closely mirrors Jev, a specialized classifier model released by startup TypeSafe AI that is built on a large language model (LLM) but designed to output probabilistic choices cheaply and rapidly, making it far more economical than full LLM inference for high-volume software automation tasks.
The core insight driving both products is that standard LLMs are too slow and expensive for many software automation use cases, especially monitoring AI agents. Developers have begun using Jev to watch over agentic AI actions in real time — checking each action against its assigned task and blocking or flagging bad behavior. A demo showed agent monitoring costing $2.94 with Jev versus $372 with a frontier LLM, a roughly 125× cost reduction. OpenAI's own agent safety measures reportedly rely on a separate 'watchdog' model at 'significant compute cost,' making a Decisions API-style solution potentially transformative for making agent oversight economically viable at scale.
Full analysis
OpenAI just shipped a cheaper way to make a model pick from a list. At Dev Day, Sam Altman announced the Decisions API: feed Luna a fixed menu of options (image categories, agent behaviors, moderation verdicts) and it returns a choice, fast and cheap, instead of generating free-form text. It copies Jev, a small classifier from startup TypeSafe AI, almost beat for beat. The demo number that's traveling: watching an AI agent's actions cost $2.94 with Jev versus $372 with a full frontier model. Roughly 125× cheaper.
What's actually being decided here: whether you keep paying full-model prices for the thousands of tiny yes/no and which-bucket calls your agent stack makes, or move them to a cheap classifier. This is easy to undo. It's an API swap on a decision gate, not a rewrite. The deadline is soft, set by whenever your inference bill on agent monitoring starts to hurt. Nothing forces your hand this quarter.
The Skeptic. TypeSafe built Jev because OpenAI's stack was too expensive for this, and OpenAI's answer is to ship the same thing and keep the margin. Fine. But copying a product is not winning a market. For Decisions API to matter, Luna has to beat a fine-tuned small open model (Llama, Mistral, Qwen) on the exact same list-picking task. That is not obvious. Anyone who cares about 125× cost cares enough to self-host a distilled classifier and skip the vendor entirely. The moat is convenience, and convenience moats leak the moment Hugging Face has a comparable checkpoint. OpenAI wants your decision-gate volume because it's huge, not because only OpenAI can serve it.
The Safety Lens. Everyone is reading this as a pricing story. It's a safety story with a price tag. Right now, if watching each agent action costs $372-equivalent, real deployments watch occasionally or not at all. Drop that to $2.94 and "monitor every action" becomes the default. That's a genuine shift. The catch: a cheap classifier that's confidently wrong on inputs it's never seen is worse than an honest "I'm not sure" from a slower model, because it fails silently and at full volume. OpenAI's own quote promises it keeps "safety protections." Unverified. A watchdog that blesses bad actions quickly is not oversight. It's theater at speed.
The Builder. Ship this Tuesday for any high-frequency gate: intent routing, action validation, content moderation, agent monitoring. The $2.94 versus $372 is not academic, it's the line between an agent-monitoring feature that pays for itself and one that eats your budget before you scale. What breaks first is your taxonomy. A fixed option set means someone has to design clean, non-overlapping categories, and dump ambiguous buckets in and you get fast, confident, wrong answers. The slower problem shows up at 90 days: your business logic moves, nobody re-curates the categories, and the classifier quietly drifts out of sync with what your agents actually do.
The Enterprise Buyer. A CTO signs for this only if the cheap path doesn't cost control. Decisions API means another workload locked to OpenAI's endpoint, priced at OpenAI's discretion, with your routing taxonomy living inside their model. Against that, a fine-tuned open classifier runs on your own hardware, in your own cloud, with audit logs you own. For regulated buyers watching AI agents in production, "we self-host the watchdog" is an easier conversation with legal than "OpenAI watches our agents for us." The convenience is real for a startup moving fast. For anyone with data-residency or audit requirements, lock-in on the safety layer is exactly the wrong place to add a dependency.
The real disagreements. The Safety Lens sees the most important oversight development of the week; the Skeptic sees a margin grab that any competent team routes around with an open model. Both can be right: agent monitoring becomes standard practice, and OpenAI captures almost none of it because the buyers who care most about cost and control self-host. The second split is Builder versus Buyer. The Builder wants to ship the OpenAI endpoint today; the Buyer warns that putting your safety classifier inside a single vendor's black box is the one dependency you'll regret.
What it hinges on. Two facts settle this. First, does Luna actually beat a fine-tuned small open model on constrained-choice tasks, or is the 125× purely a comparison against the wrong baseline (a frontier model nobody should have used for classification anyway)? Run Decisions API against a distilled Llama or Qwen, not against GPT-class inference, and see what the gap actually is. Second, does the cheap classifier hold up when real inputs fall outside its training categories, or does it fail silently? Before you wire this into anything that blocks agent actions, build a test set of weird, out-of-distribution inputs and measure how often it picks confidently wrong. If it degrades gracefully, ship it. If it fails silent, you've bought cheaper theater.
The council leans one way: the capability here is real but commodity. The 125× is a comparison against a baseline you shouldn't be running, and the cheap-classifier pattern is available to anyone with a GPU and a weekend.
Prediction: By OpenAI's next Dev Day (expected around September 2027), the cheap constrained-classifier pattern will be a commodity, and OpenAI's Decisions API will NOT be the price leader for agent-monitoring workloads, because a fine-tuned open model (Llama, Qwen, or Mistral derivative) on a leaderboard like Hugging Face's will match its constrained-choice accuracy at lower all-in cost for self-hosters.
Confidence: Medium. The pattern is reproducible, and the open ecosystem closes classification gaps fast.
Why: Constrained decoding and classification heads on top of a small LLM are a known, reproducible technique, not a trained-in moat, which is exactly why TypeSafe built Jev in a startup and OpenAI cloned it in a week. The 125× figure compares against frontier-model inference nobody should run for classification, so the real contest is against a distilled small model, and open checkpoints have repeatedly matched narrow-task accuracy at 20% of the cost within a year of a closed product defining the category. The less likely outcome is OpenAI holding a durable price-and-quality lead, which would require Luna's classifier to be meaningfully better in a way that doesn't transfer to open fine-tunes, and nothing in the announcement shows that edge.
Revisit by 2027-09-30: We're right if an open-weights classifier posts equal-or-better constrained-choice accuracy at lower self-hosted cost than Decisions API on a public benchmark. We're wrong if independent testing shows Decisions API leads both accuracy and all-in cost against the best open fine-tune on the same agent-monitoring task.
Comments