Industry story
OpenAI launches ChatGPT Work, agentic tool for non-engineers
agents guardrails model-pricing orchestration tool-use
OpenAI published the number that should give them pause: fewer than 1% of individual subscribers have touched an agent, against a billion people using ChatGPT on the web. ChatGPT Work at $20/month connects to email, Slack, Notion, and Figma and runs multi-step tasks autonomously, which is Codex generalized for everyone who doesn't write code. The problem is that "almost nobody tried it" and "almost nobody wanted it" are indistinguishable at 1% adoption, and Sam Altman calling it "delightful and safe" doesn't resolve which one it is. The unit economics don't help either: $20/month was priced against a single inference call, and a real agentic workflow burns dozens.
Full analysis
OpenAI shipped ChatGPT Work, a $20/month agent that reaches into email, calendar, Slack, Notion, and Figma and runs multi-step tasks on your behalf. It's Codex generalized for people who don't write code. The bet: agents go mainstream. The data OpenAI published alongside it says the opposite is happening so far.
What's actually being decided for a technical AI leader: not "do I buy ChatGPT Work," but "do I let an autonomous agent hold write access to my company's communication stack, and do I build my own agent layer on OpenAI's contracts or someone else's." Type 1 on the permissions question (hard to walk back once an agent has sent client emails), Type 2 on the vendor question (you can swap orchestration layers). Forcing function: none hard. This is a land-grab launch, not a deprecation. You have time to test.
The Skeptic. A billion people use ChatGPT on the web. Fewer than 1% of individual subscribers touched the agent. That is the whole story, and OpenAI published it themselves. When your own funnel converts at under one percent, the problem isn't awareness or price. Most white-collar workers do not have a task that an agent does better than copy-paste, and copy-paste never sends the wrong email. For a PM: OpenAI is betting people want a robot to run their inbox, and almost nobody has asked for that yet. The $20 price is too cheap to prove anyone will pay for ROI. The real test is an enterprise contract where someone has to show the agent saved money, and that hasn't happened. "Delightful and safe" from Sam Altman is the sound a CEO makes right before an incident.
The Safety Lens. A chatbot that says something dumb is embarrassing. An agent with write access to Slack and email that does something dumb is a breach report. Those are different risk tiers, and the launch framing collapses them. The 98% internal adoption at OpenAI is not reassurance, it's the problem: a company whose culture normalizes high-autonomy tooling is exactly the wrong place to calibrate defaults for a hospital, a law firm, or anyone under GDPR. For a PM: a confidently wrong agent that acts before anyone checks is a different failure mode than a chatbot giving a bad answer. Confirmation gates before irreversible actions, minimal-permission scoping, and an audit trail are prerequisites. Nothing in the announcement says they shipped ahead of the feature set.
The Researcher. The 98%-versus-17%-versus-1% adoption cascade tells you agentic utility is still gated on prompting skill, not access. Connecting Slack, Notion, and Figma is plumbing. The hard problem is reliability across a chain of steps where errors compound: a 95%-per-step model over ten steps completes the full task correctly about 60% of the time. For a PM: the demo works because the demo has no ambiguity, and your Tuesday inbox is nothing but ambiguity. Benchmarks measure "finished the task." Real work measures "finished it correctly, and recovered gracefully when step four went sideways." That second number is where this product lives or dies, and OpenAI hasn't published it.
The Enterprise Buyer. No CTO signs off on an agent with write permissions to corporate comms without SSO, granular scopes, per-action audit logs, data residency, and indemnification. A $20/month seat product ships with none of that by default. The buyer's question isn't "is it capable," it's "who's liable when it sends privileged material to the wrong distribution list." For a PM: the thing that wins enterprise deals is boring governance, and the thing that got announced is an autonomous inbox. Anthropic knows this, which is why Claude's enterprise posture leans hard on controllability and document fidelity. The capability gap between the two is narrow. The trust-and-controls gap is where the contract gets decided.
The Compute Pragmatist. Chat is one inference call. An agent completing a complicated task is dozens to hundreds, plus tool-call overhead, context re-injection, and retry loops when a step fails. Migrate even 5% of a billion web users' real workflows to this pattern and the load profile changes shape: longer contexts, stateful sessions, higher token burn per useful outcome. For a PM: every autonomous task costs OpenAI many times what a chat message costs, and $20/month was priced against chat economics. That math works at 1% adoption. It gets ugly precisely if the product succeeds. Watch whether OpenAI quietly caps agent runs or meters heavy users, because the unit economics force it.
Where the council splits. The Skeptic says the sub-1% number means nobody wants this, full stop. The Compute Pragmatist says the opposite risk is worse: if it does catch on, the inference bill breaks the price point. Both can't be the primary worry, and which one is right depends on a single unknown. Second split: the Researcher thinks the ceiling is model reliability over long chains, while the Enterprise Buyer thinks the ceiling is governance and liability. If reliability is the wall, better models fix it. If governance is the wall, no model release fixes it, only contracts and controls do. Those point to different bets.
What it hinges on. One belief: does multi-step reliability under real organizational entropy clear the bar where a knowledge worker trusts the agent with irreversible actions? Everything else is downstream. If it does, adoption climbs and compute becomes the binding constraint. If it doesn't, the sub-1% number holds and this is Codex's audience with a bigger tent, not a mainstream product. Before building on it: run your own eval on your workflows, ten-step tasks with real ambiguity, and measure end-to-end correctness plus recovery, not demo completion. And do not grant write scopes without a confirmation gate you control, because the default won't have one.
The Prediction. The council leans hard on the Skeptic and Researcher: the adoption gap is real, published by OpenAI, and driven by a reliability ceiling that a rebranded Codex does not clear.
Prediction: By OpenAI's next major ChatGPT product update or DevDay-style event before 2027-03-01, individual-subscriber adoption of the agentic Work/Codex product will still be in the single digits as a share of ChatGPT's consumer base, and OpenAI will not publish a clean end-to-end task-success rate for real multi-step workflows.
Confidence: Medium. The reliability ceiling is structural, but a strong model release could move it.
Why: OpenAI itself published the damning numbers: 98% internal use, 17% of org subscribers, under 1% of individuals. That spread doesn't come from lack of access to a billion-user base, it comes from agents failing silently in the middle of real tasks where errors compound across steps. A brand and a lower price don't fix compounding error, and the launch shipped no new evidence that multi-step reliability crossed the trust threshold. The opposite outcome, agents going mainstream in six months, would require the hard reliability problem to have quietly been solved, and if it had, OpenAI would be publishing that success rate instead of an adoption gap. Silence on the correctness number is the read: the audited version is smaller than the demo.
Revisit by 2027-03-01: We're right if OpenAI's own or credible third-party reporting shows individual agent adoption still in single-digit percentages and no clean multi-step task-success rate is published. We're wrong if OpenAI reports double-digit consumer adoption of the agentic product or publishes an audited end-to-end success rate above 80% on real workflows.
The tension worth watching underneath: if this call is right, the constraint on agents was never distribution. It was that the models still can't be trusted to finish a messy job without a human checking step four.
Comments