Refacto AI

Podcast episode

AI Companies Are Hiring More

ai-in-adtech cloud-costs engineering

Meta entering the cloud-compute market isn't good news for CoreWeave and Nebius — both dropped 14–17% in a single day when the news broke. Meta built GPU capacity for its own training runs; it doesn't need margin on external sales, which means it can undercut anyone who does. For ad-tech teams running inference on bidding models or creative-gen pipelines, that's real downward price pressure coming in the next 12–18 months. Meanwhile, Anthropic quietly monitored Claude Code users' proxy configurations and system prompts for months before a Reddit user surfaced it — a reminder that "closed model vendor" and "enterprise data boundary" are not the same thing.

Full analysis

The through-line this week isn't any single announcement — it's that the "AI eats jobs" thesis is getting mugged by data at the same moment that frontier labs are turning their spare GPUs into a business. For anyone building AI into an ad-tech product or an agency workflow, two numbers matter: AI-adopting firms grew headcount 10% (entry-level 12%) over two years, and the Remote Labor Index — a benchmark measuring whether AI can finish real freelance gigs a client would pay for — jumped 4× to 16.1% in eight months. Both point the same direction: capability is climbing fast, but it's landing as augmentation and bigger projects, not layoffs. Yet.

The Skeptic — The headline jobs finding is correlation wearing a causation costume, and even Ramp's own lead economist told listeners to distrust it. Firms that adopt AI aggressively are also the firms with capital, growth, and hiring budgets — of course they're hiring. The 6–12 month lag conveniently means the substitution shoe could still drop. For agencies, the Digiday piece is the tell: agencies are "out in front," clients are "stuck on governance." That gap doesn't shrink because a benchmark went up. For the PM: the study shows AI-heavy companies are hiring more, but it can't prove AI is the reason — winners were probably going to hire anyway.

The Researcher — The Remote Labor Index is the number to internalize, and it's a better instrument than the coding benchmarks everyone quotes. It scores client-acceptable deliverables — 3D models, video edits, data analysis — not synthetic pass/fail. But read the level, not just the slope: 16.1% means an 84% failure rate on real freelance work. The 4× jump is real and fast; the absolute capability is still low. That combination is exactly what you'd expect right before the curve gets interesting for creative and data categories in ad-tech. For the PM: the best model can now finish about 1 in 6 real freelance jobs to paying-client standard — up from 1 in 40 eight months ago.

The Open-Source Advocate — The quiet story is who controls the pipes. OpenAI floating 5% equity to the government, Meta standing up "Meta Compute," SpaceX billing Anthropic $1.25B a month — this is compute concentrating in a handful of hands. That's the opposite of what an open ecosystem wants. And Anthropic's Claude Code caught covertly fingerprinting users to catch distillation? Distillation — training a cheaper model on a bigger model's outputs — is precisely how open-weight competitors close the gap. Punishing it via telemetry is closed labs defending the moat. Ad-tech shops betting everything on one proprietary API should keep an open-weight fallback warm. For the PM: the companies that make the models are now also renting the computers and quietly policing who copies them.

The Compute Pragmatist — Watch the market's reaction: Meta announces a cloud business and CoreWeave drops 14%, Nebius 17%, while Meta rises 8.8%. Investors just repriced the neocloud middlemen. For ad-tech operators buying inference, this is good news short-term — more sellers (Meta, xAI/SpaceX, the neoclouds) means pricing pressure. Long-term it's concentration risk: if frontier labs with their own fleets become your compute supplier and your model supplier, you've stacked two dependencies on one vendor. The Ford anecdote is the ground truth on cost — they rehired 350 engineers after AI-only design failed, then won JD Power. Compute is cheap; getting quality out of it isn't. For the PM: the firms with the most GPUs are starting to rent them out, which should push inference prices down for now.

The Builder — The Fable 5 complaints are the operational headline nobody's pricing. A user paid $321 for a session where the model did 20% of the work and routed benign tasks to the expensive Opus tier. That's a metering-and-routing failure, and it's the kind of thing that torches your unit economics in a production ad-creative or campaign-analysis pipeline. The AWS Summit note — Agent Core's live demo breaking mid-session, roadmaps compressed to 2–3 months — says the agent infrastructure is genuinely day-one. Building autonomous ad-ops agents on this today means you own the on-call pain. For the PM: today's agents can do impressive work but also silently rack up bills doing the wrong thing.

Where they split

Three real disagreements:

  • Researcher vs. Skeptic on the RLI slope. The Researcher sees a 4×-in-8-months curve that extrapolates to creative/data disruption inside a year or two. The Skeptic sees an 84% failure rate and a benchmark that's easy to over-read. Both are looking at the same number. Whether you staff up or automate your creative/data-work depends entirely on which reading you believe.

  • Compute Pragmatist vs. Open-Source Advocate on Meta Compute. More compute sellers is a near-term win for buyers (cheaper inference). But it's also compute concentrating in the same firms that own the models — a long-term dependency trap. The cheap price today may be the bait.

  • Skeptic vs. everyone on the jobs data. If the hiring correlation is spurious, the whole "AI grows headcount" narrative agencies are using to reassure nervous clients collapses on the next recession-driven layoff cycle.

What it actually hinges on

For the ad-tech ecosystem, the decision underneath all this is build-vs.-hire in creative, data, and campaign-ops work — and it hinges on three checkable beliefs:

  1. Does the RLI curve hold in your category? Don't trust the aggregate 16.1%. Run your own eval: take 30 real deliverables your team ships (banner variants, campaign readouts, audience analyses), have Fable 5 / GPT-5.2 attempt them, and score against "would the client accept this." That's your actual number.

  2. Will inference stay cheap? Meta Compute and SpaceX are pushing prices down now. Sign nothing that assumes today's rate is permanent, and keep an open-weight (Llama/Qwen/Mistral) fallback validated so you're not hostage to one lab's routing decisions — or its telemetry.

  3. Does augmentation survive a downturn? The Ramp data was gathered in a growth phase. Model your business case on augmentation and stress-test the substitution scenario.

The council leans toward the Skeptic-plus-Researcher position: capability is climbing fast and for real, but the "hire more people" story is a growth-era artifact and the frontier still fails most real work. Augment, instrument your costs obsessively, and don't cut headcount on the strength of a correlation.

The call

Prediction: When the Center for AI Safety publishes its next Remote Labor Index update (roughly Q1 2027, on its ~6–8 month cadence), the frontier score will land between 25% and 45% — a clear rise from 16.1%, but still leaving a majority of real freelance deliverables failing client-acceptable quality.

Confidence: Medium — the 4× jump gives a strong trend line, but absolute ceiling is uncertain.

Why: The benchmark went 2.5% → 16.1% in eight months as labs shipped better agentic tool-use; the same driver continues, but each quality tier gets harder and the current 84% failure rate leaves enormous headroom that won't close in one cycle.

Revisit by 2027-03-31: We're right if the next RLI frontier score is between 25% and 45%. We're wrong if it exceeds 45% (faster disruption than augmentation thesis assumes) or comes in below 22% (the curve is already flattening).

Comments