Refacto AI

Podcast episode

41 Stats That Tell the Story of AI Right Now

agents coding-agents evals model-pricing open-weights

Nathaniel Whittemore stacked 41 stats into one argument: AI is mainstream by headcount and nowhere near mainstream by return. 52% of U.S. workers now use it on the job. 7% of global leaders report established ROI. The gap between those two numbers is the whole story.

The stats worth holding onto are the transactional ones, not the survey results. Ramp's index of 70,000-plus business customers counts actual paid subscriptions: Anthropic at 42.4%, OpenAI at 39.5%. Dollars leaving accounts, not opinions on a form. On the engineering side, AI-generated code crossed 50% of commits across 500-plus organizations in Q2, up 16 points in a single quarter. Meanwhile, 98% of teams say they worry about token costs (the per-query price of running a model) but only 64% actually meter their usage. You can't manage what you won't measure.

The shadow-AI number is the one to act on: 66% of office workers use tools they believe violate company policy. Your team is already in Claude whether you said so or not. Decide on purpose.

Full analysis

Nathaniel Whittemore stacked up 41 stats to make one claim: AI is mainstream by headcount but nowhere near mainstream by return. 52% of U.S. workers now touch it on the job. 7% of global leaders report established ROI. The rest is noise between those two numbers. For an engineering lead shipping AI into production, this is a Type 2 read: nothing here forces a decision today, but two figures should change how you instrument your stack. Nothing to reverse, plenty to measure.

The Skeptic

Half these stats are survey theater. KPMG says 76% of leaders report "meaningful business value," up 12 points in a quarter, while only 7% have established ROI. Those two numbers cancel each other out. "Meaningful value" is what an exec says when the board asks and the P&L can't answer. The Challenger job-cut data is worse: AI as the #1 stated reason for cuts, five months running, while Yale Budget Lab finds zero occupational fingerprint in the actual labor data. Whittemore nailed this one. AI is the palatable excuse for layoffs a CFO wanted anyway. For a PM: when a vendor waves an adoption percentage at you, ask what transaction it's counting. Most of these count nothing.

The Researcher

The two figures that survive scrutiny are transactional, not self-reported. Ramp's July index covers 70,000+ business customers and measures actual paid subscriptions: Anthropic at 42.4%, OpenAI at 39.5%. That's dollars leaving accounts, not opinions in a form field. DX's code-generation number is the other real one: AI-generated code crossed 50% across 500+ engineering orgs in Q2, up from 34% a quarter earlier. That's a 16-point jump in three months, measured in commits. Compare that to the EY stat where 98% claim token-cost anxiety but only 64% meter usage. When the same population won't measure the thing it claims to fear, the fear is performative and the 50% code number is not.

The Builder

The DX and OpenAI Codex stats are the ones I'd act on Monday. Over 25% of Codex users have handed the agent tasks scoped at 8+ hours of human work. That's not autocomplete. That's delegation, and it changes your review load, your CI, your blast radius when the agent ships something plausible and wrong. If half your codebase is model-written, your eval harness and your test coverage are now your actual moat, not your senior engineers' keystrokes. The shadow-AI number backs this: 66% of office workers use tools they think violate policy. Your team is already pasting code into Claude whether you sanctioned it or not. Block it and you lose the capability gap; sanction it and you inherit the data-governance risk. Pick on purpose, not by neglect.

The Compute Pragmatist

Here's the stat that should keep a platform lead up at night: median AI spend is $11.38 per employee per month, top 1% of firms hit ~$7,500. That spread is the whole story of consumption pricing. Seat-based SaaS gave you a flat, predictable bill. Agentic workloads bill by task complexity, and a single Codex job chewing through an 8-hour task burns tokens like nothing on a per-seat plan ever did. The 98%-worried, 64%-metering gap means most shops can't even see their own unit economics. If you're shipping agents to production without per-workflow token telemetry, you will get a bill you can't explain and can't attribute. Instrument cost per task before you scale the agent, not after.

The Open-Source Advocate

Notice what's absent from all 41 stats: open weights. No Llama, no Mistral, no Qwen in the Ramp index, because the index only counts paid frontier subscriptions. That's a measurement bias, not a market truth. The Pew stat that Americans believe China leads U.S. in AI by 3-to-1 sits next to a Ramp chart that literally can't see the Chinese open models eating share, and Alex Atallah makes exactly this point on the 20VC episode: Chinese open models are closing the gap while enterprises route around the closed labs. If 66% of workers run shadow AI, some meaningful slice is running a local or open model precisely to dodge the token bill and the data-governance problem in one move. The two-company race is a two-company race for paid API dollars. It is not the whole board.

Where they part ways

The Researcher trusts the Ramp spend flip as a leading indicator that coding-agent capability now drives procurement. The Open-Source Advocate says Ramp is structurally blind to the fastest-moving segment, so a 42.4/39.5 split between two closed labs tells you who's winning a shrinking definition of the market. Both can be right: Anthropic can lead paid frontier spend while open weights quietly erode what "paid frontier spend" even means.

The second split is Builder versus Compute Pragmatist on that 50% AI-generated-code number. The Builder sees capability worth shipping. The Pragmatist sees an unmetered cost curve. The decision lives in whether your token telemetry matures as fast as your agent adoption. Right now the EY data says it doesn't.

What it hinges on

Two beliefs. First: is the Anthropic lead on Ramp real capability preference, or just Claude Code being priced and packaged well for a moment? If it's capability, it persists through the next model release. If it's packaging, OpenAI's next Codex push flips it back. Second: does the 34%-to-50% code-gen jump hold, or was Q2 a novelty spike that plateaus once teams hit the review-and-cleanup wall? Verify by instrumenting your own two numbers: cost per completed agent task, and the ratio of AI-written code that survives review unedited. Those tell you more than any of the 41 stats, because they're yours.

Prediction: Ramp's spend index will still show Anthropic ahead of OpenAI on paid business subscriptions in its Q4 2025 data release, covering October through December, sometime in January 2026.

Confidence: Medium. Transactional data, and Claude Code has a real agentic-coding lead driving the spend.

Why: The Ramp flip isn't a survey blip; it's actual paid subscriptions across 70,000+ firms, and the mechanism behind it is Claude's edge on agentic coding workloads, which the DX 50%-code-gen data shows is exactly where enterprise spend is concentrating. Procurement momentum in coding tools is sticky, because teams build eval harnesses and CI around a given agent and don't rip it out mid-quarter. OpenAI could retake the lead with a strong Codex or GPT release, which is the real risk to this call, but a single model launch rarely reverses a spend-share gap this fast when switching costs are rising. The likelier path is Anthropic holds or extends the margin through year-end.

Revisit by 2026-01-31: We're right if Ramp's Q4 data keeps Anthropic above OpenAI on paid business subscription share. We're wrong if OpenAI retakes the lead in that index.

Comments