Refacto

Industry story

Holdcos Eye Token Futures Market to Monetize AI Costs

agency cost-compression inference model-pricing

Holdcos want to buy AI tokens wholesale and resell them to clients at a markup, framed as a discount versus open-market rates. It's principal media trading with a new cost line, and Ebiquity CEO Ruben Schreurs puts the forcing function plainly: agencies spent two years eating AI costs to win business and can't keep doing it. The problem is that inference prices have fallen 10 to 100x in eighteen months and keep falling, so any spread a holdco locks today is a bet that the floor holds. It won't.

Full analysis

Holdcos want to buy AI tokens wholesale, mark them up, and sell them to clients as a discount. Same play as principal media trading, where an agency buys inventory for its own book and resells it at a spread. The question for anyone building with LLMs: does bulk-buying inference actually create leverage, or is this a margin grab on a cost line that's collapsing?

Reversibility: Type 1 for the holdcos betting their procurement model on it, Type 2 for the builders who have to plumb it. What's actually being decided: not "should agencies buy tokens" but "can the reseller layer survive a market where the underlying price drops faster than the spread." Forcing function: Ruben Schreurs of Ebiquity is explicit that the subsidy ends this year. That's the clock.

The Skeptic Three things have to be true and none of them are. Token prices have to stay high enough for the spread to matter. OpenAI and Anthropic have to actually cut tier deals with resellers. And clients have to accept an opaque markup. Inference has fallen 10 to 100x in eighteen months. Capacity-constrained frontier labs don't need to discount, and every one of them is building direct-to-enterprise motion. Post-MediaLink clients audit everything. For the PM in the room: agencies are trying to profit off a cost that's cratering, which is like cornering the ice market in April.

The Compute Pragmatist The arbitrage window was 2023. Locking committed-use discounts now means betting frontier prices stabilize, and they won't. Open-weight models and hyperscaler-native inference are eating the floor. Holdcos could warehouse expensive committed capacity while a Llama or Qwen variant does 80% of the ad-copy job for a tenth of the cost, running on rented GPUs the client already pays for. Worse, the labs can tier-price to recapture any margin a reseller extracts. For the PM: buying a year of tokens in advance is smart only if the price goes up. It's going down.

The Safety Lens A reseller who profits per token has a structural reason to push more inference, even where a boring deterministic workflow would be cheaper and more auditable. That's backwards from responsible use. The deeper exposure is provenance. Once tokens flow through a holdco layer, the client can't easily see which model, which fine-tune, which safety config touched their work. For a finance or healthcare client under real regulatory load, that opacity is compliance exposure, not just a trust problem. For the PM: if you can't name which model wrote the thing, you can't defend it when a regulator asks.

The Enterprise Buyer This is where the model breaks fastest. A CTO signing a services contract now expects token-level logs, per-client cost attribution, and pass-through pricing, because the whole industry learned that lesson from principal media. The pitch is "discount versus open market." Prove it. Show me the rate card, show me the direct API price, show me the delta. Most current LLM billing wasn't built for multi-tenant resale, so the holdco either builds real reconciliation infrastructure or the numbers don't survive an audit. Enterprise buyers who already have their own OpenAI or Anthropic enterprise agreement have zero reason to buy through an intermediary.

Where they part ways

The Skeptic and Compute Pragmatist agree the spread evaporates, but for different reasons: the Skeptic thinks clients won't tolerate the markup, the Pragmatist thinks the price drop kills it before clients even notice. Those aren't the same bet. If clients are asleep, a shrinking spread can still be sold as a discount for a while.

The Safety Lens sees a durable problem the others treat as transient. Even if the economics fail, the provenance gap is real the moment one token flows through a reseller. That's the piece a builder can't paper over with a reconciliation job.

What it hinges on

One belief: can holdco volume actually move wholesale token pricing enough to fund a spread the client won't claw back? Nobody has the consumption numbers to prove it. The Ebiquity line about two years of subsidy is one operator's anecdote, not a dataset. Before any agency-side team wires this in, get the direct enterprise API price in writing, then demand the holdco show the negotiated wholesale rate against it. If they won't, there's no discount to verify, only a markup to hide.

Prediction: By the end of Q1 2027 earnings calls, with Omnicom, WPP, and Publicis all reporting in the February to March 2027 window, no major holdco will report a material, separately disclosed revenue line from reselling AI tokens to clients at a markup.

Confidence: Medium. Deflating token prices and audit-wary clients gut the spread.

Why: Inference prices have dropped 10 to 100x in eighteen months and every observable signal points down, so any wholesale spread a holdco locks today shrinks under it while open-weight and hyperscaler-native inference undercut the floor. Frontier labs are capacity-constrained and building direct-to-enterprise, giving them no reason to hand reseller margin to an intermediary, and post-MediaLink clients demand pass-through and token-level logs that current LLM billing can't cleanly produce. For this to become a real, disclosed revenue line by early 2027, all three would have to break the holdcos' way inside a year, and the trend runs against each. The likelier outcome is quiet pilots and consulting-flavored bundling, not a standalone token-arbitrage business anyone puts on a slide.

Revisit by 2027-03-31: We're right if no holdco breaks out token resale as a disclosed revenue line on its FY2026 or Q1 2027 results. We're wrong if any of Omnicom, WPP, Publicis, or Dentsu reports token resale as a named, material contributor.

The tell will be in the audit clauses, not the press releases. If clients start writing token pass-through into their contracts, the spread was never going to hold.

Comments