Refacto

Industry story

Stealth crawler bot traffic surged 597% in 2025, threatening publishers

brand-safety measurement programmatic publisher-economics

AI scraper traffic — bots that silently pull content from publisher websites without identifying themselves — grew 597% from January to December 2025, according to cybersecurity firm HUMAN Security. Overall AI-driven web traffic nearly tripled year-over-year, and scraping attacks now affect nearly 20% of median site traffic, almost double 2022 levels. Cloudflare data shows more than half of all web traffic is now bot-based, making it increasingly difficult for publishers to detect unauthorized content extraction or attribute it to specific AI companies or data resellers.

Publishers are caught in the dark because stealth crawlers mask themselves as ordinary browsers or use residential IP addresses, bypassing robots.txt rules (the standard files websites use to tell crawlers what they may or may not access). One publishing executive, Lindsay Van Kirk of People Inc., described going from blocking roughly 2,100 user agents to over 30,000 after adopting a block-all-bots strategy — equivalent to tens of millions of scrape attempts per day. Media analyst Matthew Scott Goldstein estimates the broader scraper economy has grown into at least a $1 billion industry, with some executives suggesting it is actually multi-billion dollar in scale.

Full analysis

Bots now outnumber people on the open web. Cloudflare says more than half of all web traffic is bot-based, and HUMAN Security clocked stealth scraper traffic up 597% across 2025. For an ad-tech operator, this is a question dressed as a security story: what fraction of the audience you're selling is real, and what do you do when a growing chunk of your traffic is machines pulling your content to train someone else's model?

Reversibility: Type 1 for the open web's structure. Once buyers decide open-web pageviews are contaminated, that trust doesn't come back cheaply. Type 2 for any individual publisher's blocking policy, which you can dial up or down.

What's actually being decided: Not "should I block bots." It's whether open-web publishers keep monetizing inventory the way they always have, or restructure toward authenticated, bot-resistant environments before the measurement credibility erodes underneath them.

Forcing function: Q1/Q2 2026 budget planning, and the next round of measurement-vendor reconciliation when human reach and reported pageviews start to diverge.


The Market Analyst. Follow who gets paid. HUMAN Security and Cloudflare sell the shovels in this gold rush, and every scary stat they publish is also a sales deck. That doesn't make the trend fake, it makes the framing self-interested. The real market move is quieter: if bot-polluted pageviews inflate the denominator, then open-web CPMs look stable while actual human reach shrinks. Buyers eventually notice. That pushes brand budgets toward authenticated inventory: Reddit, the NYT, premium direct deals, and CTV, where a logged-in human is a logged-in human. In plain terms: the ad market is slowly learning that a chunk of "people" it's been buying were never people.

The Skeptic. A 597% growth number from a company selling bot mitigation is a marketing artifact until proven otherwise. The stat also lumps stealth scrapers in with legitimate AI traffic: indexing, analytics, SEO crawlers. If this were torching real ad revenue, DoubleVerify and IAS would be flagging invalid-traffic spikes in their earnings commentary, and they haven't, not dramatically. The $1 billion scraper economy is analyst Matthew Scott Goldstein's estimate, not audited anything. Here's the distinction that actually matters: scrapers steal your content, but they mostly don't load your ads. Training theft and impression fraud are different problems, and this story blends them. Plain version: your content is being taken, but that's not the same as your ad money being stolen.

The Operator. Lindsay Van Kirk of People Inc. went from blocking 2,100 user agents to more than 30,000. That's not an IT ticket anymore, that's a standing ops function with a headcount cost nobody budgeted. The 90-day trap: ad fill and pageview forecasts run off inflated numbers, so CPMs look fine while real reach quietly bleeds. Your revenue ops team and your ad ops team need to be in the same meeting, and right now they aren't. Block too aggressively and you'll catch Googlebot or a paying partner's crawler in the net, and then you've torched real revenue to stop a scrape that never cost you a cent. Plain version: the cure can hurt more than the disease if you swing blind.

The CFO. Two different line items masquerading as one. Content theft is a licensing question: OpenAI, Anthropic, and the rest are taking inventory I might otherwise sell them. That argues for a deal, not a firewall. Impression integrity is a revenue question: if bots inflate my traffic, I'm forecasting off fiction and my CPMs are a mirage. I'll spend real money on Cloudflare and HUMAN before I'll spend it chasing a content-licensing fantasy, because most mid-tier publishers have no leverage to license anything. The uncomfortable math: blocking costs money now, and the payback is a number I can't see, prevented erosion. That's a hard check to write.


Where the council splits.

The Skeptic versus the Market Analyst on scale. Is this a vendor-inflated scare, or a genuine erosion of what programmatic can credibly sell? Both can't be right about magnitude.

The CFO versus the Operator on what problem you're solving. Content theft wants a licensing conversation. Impression pollution wants a real-time blocking layer. Conflate them and you'll spend against the wrong one.

The deepest tension: does aggressive blocking protect your business or shrink it? Every bot you block is content you keep, but the arms race never ends, and you're paying rent to Cloudflare and HUMAN forever to fight it.


What this hinges on. Two beliefs. First, whether bot-inflated pageviews are materially distorting human reach in a way buyers will price in. Second, whether stealth scrapers touch your ad revenue at all, or just your content. Most of the panic conflates these.

The council leans one way: the measurement-credibility risk is real and slow, the impression-fraud risk is probably overstated, and the content-theft problem is real but only actionable if you have licensing leverage, which most open-web publishers don't. So the honest move is defensive on measurement (reconcile server-side impressions against human-verified reach before your forecasts drift) and skeptical on the headline number.

Before committing budget: pull your own logs. Compare human-verified sessions against reported pageviews. If they've diverged more than a few points over the past year, the Operator is right and this is a Q1 problem. If they haven't, the Skeptic is right and you're buying a vendor's fear.


Prediction: By the end of Q1 2026 earnings calls (late April 2026), neither DoubleVerify nor Integral Ad Science will report a material spike in invalid-traffic rates attributable to AI scrapers, confirming that stealth crawlers are hitting content, not ad revenue.

Confidence: Medium. Scrapers pull HTML; they don't render and load ad impressions.

Why: Stealth crawlers exist to harvest text for training, and they generally don't execute the JavaScript that fires an ad impression, so they largely bypass the very inventory that DV and IAS measure. If this were inflating measured invalid traffic, the verification vendors, whose entire business is counting bad impressions, would be the first to flag it and the loudest, because it sells more product. The absence of that flag so far, despite a 597% scraper surge, tells you the two problems live in different places. The opposite outcome, a sudden IVT spike tied to AI crawlers, would require scrapers to start rendering full pages at scale, which is expensive and pointless for their actual goal.

Revisit by 2026-04-30: We're right if DV and IAS Q1 earnings commentary shows no material AI-scraper-driven IVT increase. We're wrong if either names AI crawlers as a meaningful new source of invalid traffic in reported metrics.

Comments