Podcast episode
AI Companies Are Hiring More
ai-in-adtech cloud-costs engineering
TL;DR
This episode covers three distinct AI-ecosystem topics: (1) new labor-market data showing AI-adopting companies are actually growing headcount faster, not cutting it; (2) OpenAI's reported proposal to give the US government a 5% equity stake; and (3) Meta's plans to monetize excess AI compute via a new cloud business. Worth listening for the Ramp/Revelio Labs jobs data and the Claude Code surveillance rollback.
What was covered
-
OpenAI equity-to-government proposal: OpenAI reportedly floated giving the US government a 5% stake (~$42B at current valuation) to be placed in a sovereign wealth fund modeled on the Alaskan Permanent Fund. OpenAI also proposed that all leading AI developers — Anthropic, Google, Meta — contribute 5% equity. Talks described as "conceptual" and potentially requiring an act of Congress.
-
Meta Compute cloud business: Bloomberg reports Meta is developing a cloud business called "Meta Compute," led by head of infrastructure Santosh Jennerdhan alongside Daniel Gross (Meta Superintelligence Labs) and Dina Powell McCormick. Two models under consideration: (a) hosted model access similar to AWS Bedrock; (b) raw compute sales similar to neoclouds like CoreWeave. Meta stock rose 8.8% on the news; CoreWeave fell 14%, Nebius fell 17%.
-
SpaceX/xAI compute deals context: Cited as the template — SpaceX signed a $1.25B/month compute deal with Anthropic in early May, followed by deals with Google and Reflection AI. Compute sales now estimated to be SpaceX's primary revenue driver ahead of Starlink.
-
Claude Code surveillance rollback: Anthropic had embedded code in Claude Code since early April that detected whether users had a proxy enabled and covertly transmitted timezone/metadata back to Anthropic to identify possible Chinese-lab users engaged in model distillation (copying a model's outputs to train a competing model). After a Reddit user exposed it, Claude Code developer Tariq confirmed the rollback, describing it as an "experiment" to prevent account abuse and distillation.
-
Claude 4 (Fable 5) re-release user reactions: After Anthropic reset usage limits, users reported strong performance on complex coding and agentic orchestration tasks, though significant complaints emerged about benign coding tasks being routed to Opus (one user reported Fable doing only 20% of work, paying $321 for a session where Fable "refused to do the work"). Host notes this is day-one territory for long-running agentic workflows.
-
Center for AI Safety Remote Labor Index: New benchmark measuring AI's ability to complete freelance tasks (3D modeling, graphic design, video editing, data analysis, web development) at a quality level a paying client would accept. Claude 4/Fable 5 scored 16.1% vs. GPT-5 at 6.3% and Opus at 8.3%. GPT-5.2 (the top performer when the benchmark launched ~8 months ago) scored only 2.5% — representing more than a 4× increase at the frontier in under eight months.
-
Ramp/Revelio Labs AI-adoption vs. hiring study (21,000 US firms): Companies with high AI adoption grew headcount 10% on average over two years; low-adoption firms were flat. Entry-level headcount grew at 12% — faster than overall. Threshold for "high adoption" was modest: ~$30/employee/month in early phases. Headcount growth lagged AI adoption by 6–12 months. Box CEO Aaron Levy's parallel survey of 1,600+ mid/large companies found 58% expected headcount to rise over three years, climbing to 79% among the most mature AI adopters.
Notable claims & predictions
-
Center for AI Safety: "The frontier has more than quadrupled in under eight months — a concrete signal of how quickly economically capable AI agents are advancing." (On the Remote Labor Index going from 2.5% to 16.1%.)
-
OpenAI Chief Economist Ronnie Chatterjee (at European Central Bank event): "Just because a task is exposed to AI doesn't mean it's going to substitute for that… those jobs [software developers] shrinking as AI capabilities increased — that really hasn't happened to the same extent people were predicting."
-
Ramp Lead Economist Eric Kharazian: "This is our first evidence that high AI-adopting firms are hiring different kinds of employees. We believe they are selecting for a new set of skills — specifically people who know how to use AI and use it well." (Notably cautioned listeners to be skeptical of the causal inference.)
-
Box CEO Aaron Levy: "If a company can get more customers because they use AI in sales… they hire more salespeople, not fewer. If you can build way more software than before, you end up hiring more engineers because the project gets bigger and you take on more."
-
Ford VP Charles Poon: "Mistakenly, we thought that just by introducing AI and ingesting the design requirements that we had, we could produce a high-quality product." Ford rehired 350 veteran engineers to retrain AI systems after quality failures — and subsequently topped the JD Power initial quality survey.
Why this matters for AI operators
-
Inference economics are bifurcating at the frontier: Meta Compute entering the market as a hyperscale compute seller (on top of SpaceX/xAI's $1.25B/month Anthropic deal) signals that frontier labs with large GPU fleets will increasingly compete directly with neoclouds like CoreWeave. Operators buying inference or raw compute should expect pricing pressure and new supplier options — but also concentration risk if Big Tech dominates supply.
-
The Ramp/Revelio data materially complicates workforce planning assumptions: The finding that AI-adopting firms grow headcount 10–12% faster (with a 6–12 month lag) — at a relatively modest $30/employee/month spend threshold — suggests the "AI replaces headcount" assumption underlying many enterprise ROI models is empirically weak, at least in this phase. AI operators and enterprise buyers should recalibrate internal business cases accordingly.
-
Claude Code's surveillance rollback is a supply-chain trust signal: Anthropic embedding covert system-prompt modifications to detect Chinese-lab users is an early example of frontier labs using their model-deployment position for competitive intelligence. Enterprises deploying Claude Code in regulated environments (finance, healthcare, government) should treat this as a reminder that model providers retain significant telemetry visibility — and that "rolled back" features can be re-introduced.
-
The Remote Labor Index (16.1% at frontier) is a more operationally honest benchmark than SWE-bench: It measures client-acceptable quality on real freelance deliverables, not pass/fail on synthetic tasks. The 4× improvement in 8 months is the sharpest near-term signal yet for operators making build-vs.-hire decisions in creative, coding, and data-work categories — though the 84% failure rate means human-in-the-loop workflows remain economically necessary for most production use cases.
Full analysis
The through-line this week isn't any single announcement — it's that the "AI eats jobs" thesis is getting mugged by data at the same moment that frontier labs are turning their spare GPUs into a business. For anyone building AI into an ad-tech product or an agency workflow, two numbers matter: AI-adopting firms grew headcount 10% (entry-level 12%) over two years, and the Remote Labor Index — a benchmark measuring whether AI can finish real freelance gigs a client would pay for — jumped 4× to 16.1% in eight months. Both point the same direction: capability is climbing fast, but it's landing as augmentation and bigger projects, not layoffs. Yet.
The Skeptic — The headline jobs finding is correlation wearing a causation costume, and even Ramp's own lead economist told listeners to distrust it. Firms that adopt AI aggressively are also the firms with capital, growth, and hiring budgets — of course they're hiring. The 6–12 month lag conveniently means the substitution shoe could still drop. For agencies, the Digiday piece is the tell: agencies are "out in front," clients are "stuck on governance." That gap doesn't shrink because a benchmark went up. For the PM: the study shows AI-heavy companies are hiring more, but it can't prove AI is the reason — winners were probably going to hire anyway.
The Researcher — The Remote Labor Index is the number to internalize, and it's a better instrument than the coding benchmarks everyone quotes. It scores client-acceptable deliverables — 3D models, video edits, data analysis — not synthetic pass/fail. But read the level, not just the slope: 16.1% means an 84% failure rate on real freelance work. The 4× jump is real and fast; the absolute capability is still low. That combination is exactly what you'd expect right before the curve gets interesting for creative and data categories in ad-tech. For the PM: the best model can now finish about 1 in 6 real freelance jobs to paying-client standard — up from 1 in 40 eight months ago.
The Open-Source Advocate — The quiet story is who controls the pipes. OpenAI floating 5% equity to the government, Meta standing up "Meta Compute," SpaceX billing Anthropic $1.25B a month — this is compute concentrating in a handful of hands. That's the opposite of what an open ecosystem wants. And Anthropic's Claude Code caught covertly fingerprinting users to catch distillation? Distillation — training a cheaper model on a bigger model's outputs — is precisely how open-weight competitors close the gap. Punishing it via telemetry is closed labs defending the moat. Ad-tech shops betting everything on one proprietary API should keep an open-weight fallback warm. For the PM: the companies that make the models are now also renting the computers and quietly policing who copies them.
The Compute Pragmatist — Watch the market's reaction: Meta announces a cloud business and CoreWeave drops 14%, Nebius 17%, while Meta rises 8.8%. Investors just repriced the neocloud middlemen. For ad-tech operators buying inference, this is good news short-term — more sellers (Meta, xAI/SpaceX, the neoclouds) means pricing pressure. Long-term it's concentration risk: if frontier labs with their own fleets become your compute supplier and your model supplier, you've stacked two dependencies on one vendor. The Ford anecdote is the ground truth on cost — they rehired 350 engineers after AI-only design failed, then won JD Power. Compute is cheap; getting quality out of it isn't. For the PM: the firms with the most GPUs are starting to rent them out, which should push inference prices down for now.
The Builder — The Fable 5 complaints are the operational headline nobody's pricing. A user paid $321 for a session where the model did 20% of the work and routed benign tasks to the expensive Opus tier. That's a metering-and-routing failure, and it's the kind of thing that torches your unit economics in a production ad-creative or campaign-analysis pipeline. The AWS Summit note — Agent Core's live demo breaking mid-session, roadmaps compressed to 2–3 months — says the agent infrastructure is genuinely day-one. Building autonomous ad-ops agents on this today means you own the on-call pain. For the PM: today's agents can do impressive work but also silently rack up bills doing the wrong thing.
Where they split
Three real disagreements:
-
Researcher vs. Skeptic on the RLI slope. The Researcher sees a 4×-in-8-months curve that extrapolates to creative/data disruption inside a year or two. The Skeptic sees an 84% failure rate and a benchmark that's easy to over-read. Both are looking at the same number. Whether you staff up or automate your creative/data-work depends entirely on which reading you believe.
-
Compute Pragmatist vs. Open-Source Advocate on Meta Compute. More compute sellers is a near-term win for buyers (cheaper inference). But it's also compute concentrating in the same firms that own the models — a long-term dependency trap. The cheap price today may be the bait.
-
Skeptic vs. everyone on the jobs data. If the hiring correlation is spurious, the whole "AI grows headcount" narrative agencies are using to reassure nervous clients collapses on the next recession-driven layoff cycle.
What it actually hinges on
For the ad-tech ecosystem, the decision underneath all this is build-vs.-hire in creative, data, and campaign-ops work — and it hinges on three checkable beliefs:
-
Does the RLI curve hold in your category? Don't trust the aggregate 16.1%. Run your own eval: take 30 real deliverables your team ships (banner variants, campaign readouts, audience analyses), have Fable 5 / GPT-5.2 attempt them, and score against "would the client accept this." That's your actual number.
-
Will inference stay cheap? Meta Compute and SpaceX are pushing prices down now. Sign nothing that assumes today's rate is permanent, and keep an open-weight (Llama/Qwen/Mistral) fallback validated so you're not hostage to one lab's routing decisions — or its telemetry.
-
Does augmentation survive a downturn? The Ramp data was gathered in a growth phase. Model your business case on augmentation and stress-test the substitution scenario.
The council leans toward the Skeptic-plus-Researcher position: capability is climbing fast and for real, but the "hire more people" story is a growth-era artifact and the frontier still fails most real work. Augment, instrument your costs obsessively, and don't cut headcount on the strength of a correlation.
The call
Prediction: When the Center for AI Safety publishes its next Remote Labor Index update (roughly Q1 2027, on its ~6–8 month cadence), the frontier score will land between 25% and 45% — a clear rise from 16.1%, but still leaving a majority of real freelance deliverables failing client-acceptable quality.
Confidence: Medium — the 4× jump gives a strong trend line, but absolute ceiling is uncertain.
Why: The benchmark went 2.5% → 16.1% in eight months as labs shipped better agentic tool-use; the same driver continues, but each quality tier gets harder and the current 84% failure rate leaves enormous headroom that won't close in one cycle.
Revisit by 2027-03-31: We're right if the next RLI frontier score is between 25% and 45%. We're wrong if it exceeds 45% (faster disruption than augmentation thesis assumes) or comes in below 22% (the curve is already flattening).
Comments