Industry story
Holdcos Explore Buying AI Tokens in Bulk, Reselling at Margin
agency ai-in-adtech cost-compression
Advertising holding companies (large agency groups like WPP, Publicis, Omnicom) are exploring a model where they pre-purchase AI compute 'tokens' — the units of processing that power AI tools — in bulk at negotiated rates, then resell them to clients at a margin, mirroring how they have long traded media inventory. The rationale: agencies have spent two years absorbing AI infrastructure costs on their own balance sheets, and those costs are no longer sustainable as AI usage scales. Token prices are also widely believed to be artificially subsidized today, creating a hedge incentive to lock in volume now before prices rise. Some holdcos are already baking token costs into so-called 'principal media' deals — arrangements where the agency buys inventory as a principal rather than an agent, taking on risk in exchange for a margin.
Full analysis
Your draft
The holding companies want to buy AI compute in bulk and resell it to clients at a markup, the same way they've traded media inventory for decades. That's the pitch. What's actually being decided is who eats the cost of AI as it scales, now that two years of agencies quietly absorbing it on their own books has stopped being tenable.
This is a Type 2 decision, easy to reverse. Nobody has committed to a three-year token contract yet. It's a trial balloon floated through a Digiday piece, and the forcing function is Q2 and Q3 earnings, where holdco CFOs need a margin story that isn't "AI is eating our profit." The question for you as an operator: is the arbitrage desk a real business line worth building toward, or a distraction from where the actual money is?
In plain terms for a non-specialist: agencies are asking whether they can become a middleman for AI computing power, buying it cheap in bulk and selling it on at a profit, betting the price goes up. This story matters to anyone who sells services to advertisers, because it's really about who pays for AI.
The Market Analyst. WPP has the most urgency here, and not for good reasons. Margin pressure, a chair transition, and the S4 Capital wreckage as a warning about betting the firm on an AI narrative. Publicis is the one that could actually make token-pool economics work, because Epsilon already gives it a data-infrastructure spine to bolt compute onto. Omnicom and Dentsu are watching. The tell is that this arrives as an exploration piece, not a press release, which means the client conversation about markup hasn't been won. Expect it on Q2 calls as a margin-recovery line, and expect analysts to ask the obvious question: do clients accept the fee, or do they just call OpenAI directly?
For the generalist: the agencies most desperate to find new profit are the ones floating this hardest, which is itself a warning.
The Skeptic. The whole model rests on token prices rising. They aren't. Inference costs have fallen more than 90% in eighteen months, and there's no mechanism to reverse that. So the hedge thesis, lock in volume before prices spike, is betting against the clearest trend in the industry. Robert Webster nailed the other trap in the quote: get good at prompting or agentic orchestration, need a fraction of the tokens you committed to, and now you're sitting on inventory you have to dump or write off. You improve your own product and it blows up your arbitrage book. That's a business that punishes you for being good at your job.
For the generalist: they'd be locking in a high price for something that keeps getting cheaper.
The Operator. Say the strategy is right. Tuesday morning, procurement asks who owns the token liability on the balance sheet and how it gets allocated to client P&Ls. Nobody has an answer. The first thing that breaks is the client contract, because master service agreements were never written for compute resale at a margin, so legal redlines every deal and the whole thing slows to a crawl. At ninety days you've created a new internal SKU, "AI compute," that account teams have no pricing muscle for, and CFOs start demanding utilization reports on the pre-bought pools. Now you need a token-ops function that doesn't exist and nobody budgeted for.
For the generalist: even if the idea is sound, agencies aren't built to run it.
The Customer / End User. Put yourself in the client's chair. Their CMO's team already knows AI compute is cheap and getting cheaper, because the vendors email them about it. So when the agency lands a line item marked "AI compute plus margin," the procurement lead asks why they can't just buy it direct. The honest answer is they can. Principal media already draws fire in the EU and UK for opacity, clients hate not seeing the real price, and now you want to layer opaque compute markup into the same structure. That invites the same audit. The clients who tolerate this are the ones too small to negotiate their own API deals, which is not where the holdco money is.
For the generalist: big advertisers can already buy this stuff themselves, and they know it.
The CFO. The appeal is obvious and it's the trap. Two years of AI costs sitting on the P&L with no offsetting revenue, and here's a structure that turns a cost center into a margin line. But look at the real cost. You commit to three-year volumes on a deflating asset, you build a token-ops and reporting function from scratch, and you carry write-off risk if usage comes in under forecast. The opportunity cost is worse: every dollar and every smart person aimed at the arbitrage desk is a dollar not aimed at orchestration IP, which is the thing clients would actually pay a durable premium for. This pays back only if prices rise. They won't.
Three real disagreements sit under this.
First, the direction of token prices. The Market Analyst grants the hedge logic enough respect to expect it on earnings calls. The Skeptic and the CFO say the deflation trend makes the entire bet backwards. Everything downstream turns on this one belief.
Second, whether the labs allow it at all. OpenAI and Anthropic both sell direct to enterprise. A resale layer that inserts an agency between them and a Fortune 500 client compresses their own relationship, and API terms can tighten to kill it. The Operator is busy solving internal plumbing; the labs may make the plumbing moot.
Third, self-cannibalization. Webster's point is the deepest one here. The better your engineers get at orchestration, the fewer tokens you need, the more committed inventory you're stuck with. The one capability worth building actively destroys the arbitrage book.
What this actually hinges on: do token prices rise or fall, and do the labs tolerate a reseller. Both point the same way. Prices are falling fast, and the labs have every reason to defend their direct enterprise motion. The council leans hard against the token-futures desk as a strategy. It's a distress signal dressed as innovation, a way to tell clients "you're going to pay for this" without saying it plainly.
The move worth de-risking is the opposite of arbitrage. The durable margin is in orchestration and prompt engineering, the IP that makes a given outcome cost fewer tokens. That compounds, it doesn't expire, and it doesn't require betting against physics. Before anyone signs a three-year volume commitment, model the write-off scenario Webster describes, because your own team getting better is the base case, not the risk case.
Prediction: No major holdco (WPP, Omnicom, Publicis, Dentsu, Havas) will announce a signed multi-year fixed-rate AI token purchase-and-resale commitment on or before its Q4 2026 earnings call in early 2027.
Confidence: Medium. Deflating token prices and lab resistance make the commitment irrational to sign.
Why: The entire model needs token prices to rise, but inference costs have fallen more than 90% in eighteen months with no mechanism to reverse, so locking in three-year volumes means overpaying on a deflating asset. On top of that, OpenAI and Anthropic both sell direct to enterprise and have every incentive to tighten API terms against a reseller that compresses their client relationships. The opposite outcome, a signed multi-year deal, would require a holdco to bet against the clearest cost trend in the industry while the labs stand aside, and Webster's own write-off trap means the agency's improving orchestration makes the committed volume a liability. Floating this through a Digiday exploration piece rather than a press release is exactly what firms do when the client conversation about markup hasn't been won.
Revisit by 2027-02-28: We're right if no holdco has publicly announced a signed multi-year fixed-rate token resale commitment by its Q4 2026 earnings call. We're wrong if any of the five names a specific multi-year token purchase-and-resale arrangement in an earnings call, investor deck, or press release before then.
Comments