Refacto

Industry story

AI Buying Agents Seen as Confidently Wrong 'Regularly,' Experts Warn

ai-in-adtech dsp measurement programmatic

When asked whether they had personally observed an AI media-buying agent produce a confident but incorrect outcome, multiple experts confirmed it happens frequently. Robert Webster of TAUMS said it occurs 'regularly,' including with his own products. Brian O'Kelley said anyone who hasn't seen it 'probably haven't used one.' Mike Follett of Lumen Research offered the sharpest framing: AI just makes bad recommendations 'quicker and prettier.' The concern is that autonomous agents — software systems that can independently decide where to place ads and how much to spend — amplify existing measurement problems by acting on flawed inputs at machine speed, making errors harder to catch before money is wasted.

Full analysis

Eight experts got asked a simple question: have you ever watched an AI media-buying agent be confidently, expensively wrong? The answer, near-unanimous, was yes. Robert Webster of TAUMS said it happens "regularly," including with his own products. Brian O'Kelley, who runs Scope3 and built AppNexus, said anyone who hasn't seen it "probably hasn't used one." Mike Follett of Lumen Research put it best: AI just makes bad recommendations "quicker and prettier."

Here's what's actually being decided. Not whether agents work. Whether operators keep a human in the approval loop, or let the machine spend at machine speed. That choice is easy to undo on paper and hard to undo in practice, because once you've cut the trafficking QA headcount that used to catch the errors, you can't summon those people back in a hurry. The deadline is set by the sales cycle: vendors are pitching "autonomous" as a feature this budget season, and someone will buy the full version to save money on staff.

The Council

The Market Analyst. Follow where the money reprices. If agents need clean, auditable inputs to act safely, then the value shifts to whoever owns the verified signal, not whoever owns the smartest agent. That favors the measurement and verification names, DoubleVerify, IAS, iSpot, and it gives O'Kelley's Scope3 a reason to exist beyond carbon data. It hurts pure autonomous-agent startups that have a clever optimizer and no ground-truth story. The gold rush is the company selling the robot a map it can trust, not the robot buyer itself. Watch which measurement vendor rebrands as "agent-ready" first. That's the moment the category has picked its lane.

The Skeptic. Webster and O'Kelley both sell things next to the problem they're diagnosing. Not disqualifying, but say it out loud. "Confidently wrong regularly" describes human traders too. The question is error rate and dollar magnitude, and nobody in this story has that denominator. Follett's line is quotable and not causal. Bad recommendations faster is only a disaster if humans have actually left the approval loop, and most enterprise deployments haven't. Much of what's called an "AI media-buying agent" today is a chatbot bolted onto a DSP. The threat is getting priced before the capability exists at scale. The fear is real. The evidence is anecdotes from vendors.

The Operator. The break isn't the agent hallucinating a CPM. It's the audit trail. When a human trader makes a bad call, there's a Slack message, a rationale, a timestamp. When an agent burns $200K on the wrong segment, the log says "optimized toward goal." That's useless to a client demanding to know what happened. Trafficking QA breaks first, because nobody staffed for catching errors at machine speed. And there's a trap in the wiring: operators trust a system's output more than a human's, especially under deadline. So the wrong number gets waved through faster. Expect emergency budget holds and manual override rules as the duct-tape fix.

The Customer / End User. Put yourself on the brand side. A CMO signed off on autonomous buying to cut cost and move faster. The first time an agent dumps six figures into garbage inventory and the agency can't explain why, that CMO stops trusting the whole category. Advertisers aren't asking for autonomy. They're asking for performance they can defend to their CFO. "The AI decided" is not an answer that survives a QBR. The vendors selling full autonomy are solving a problem the buyer doesn't have, while ignoring the one the buyer actually loses sleep over: accountability.

Where they disagree

Two real splits. First, the Market Analyst thinks clean data is the moat and the winning stack is confident agent plus auditable signal layer. The Skeptic thinks that's a tidy story that ignores how messy real incentives are, and that the "agent" mostly isn't autonomous yet anyway. One says build the infrastructure now; the other says you're pricing a capability that doesn't ship at scale.

Second, the Operator and the Customer agree the failure is accountability, not accuracy, but they squeeze different people. The Operator says ad ops eats it. The Customer says the agency loses the account. Same broken audit trail, different body count.

What it hinges on

The decision turns on one belief: is the human still in the loop when the money moves? If yes, "confidently wrong" is an annoyance, caught before spend, and the whole panic is oversold. If no, it's a real liability, and the missing piece is the explanation, the log that says why the agent did what it did.

The council leans toward the Skeptic on the near term and the Analyst on the medium term. Right now, most of this is a chatbot on a DSP with a human still clicking approve, so the catastrophe is theoretical. But the pressure to remove that human is commercial and constant, because that's where the cost savings live. When it happens, the vendor with an auditable signal layer wins the re-buy, and the pure optimizer play doesn't.

Before committing budget to any "autonomous" pitch, verify one thing: can the agent produce a human-readable reason for every spend decision above a threshold you set? If the log only says "optimized toward goal," you don't have an agent you can defend. You have a liability that scales.

Prediction: By the 2027 upfront season, at least one major measurement or verification vendor (DoubleVerify, IAS, or iSpot) will launch and market a productized "agent-ready" or agent-verification data offering aimed at feeding autonomous buying tools auditable inputs.

Confidence: Medium. The commercial logic is strong, but timing depends on how fast buyers demand it.

Why: This story shows eight practitioners agreeing that agents act confidently on flawed inputs, which turns "verified, auditable signal" into the thing that makes autonomy safe to deploy. Measurement vendors already sell exactly that raw material and are hunting for growth stories as their core verification businesses mature. Repositioning existing signal as "agent-ready" costs them almost nothing and lets them ride the agent hype instead of being disintermediated by it. The opposite outcome, all three staying quiet, would require them to watch a new buying layer form on top of their data without planting a flag, which cuts against how aggressively this category chases every adjacent narrative.

Revisit by 2027-06-30: We're right if DoubleVerify, IAS, or iSpot has publicly launched and named a product or data offering specifically positioned for AI buying agents or agent verification by then. We're wrong if none of the three has shipped or marketed such an offering by that date.

Also covered this issue

Comments