Refacto

Industry story

Stealth crawler bot traffic surged 597% in 2025, threatening publishers

brand-safety measurement programmatic publisher-economics

AI scraper traffic — bots that silently pull content from publisher websites without identifying themselves — grew 597% from January to December 2025, according to cybersecurity firm HUMAN Security. Overall AI-driven web traffic nearly tripled year-over-year, and scraping attacks now affect nearly 20% of median site traffic, almost double 2022 levels. Cloudflare data shows more than half of all web traffic is now bot-based, making it increasingly difficult for publishers to detect unauthorized content extraction or attribute it to specific AI companies or data resellers.

Publishers are caught in the dark because stealth crawlers mask themselves as ordinary browsers or use residential IP addresses, bypassing robots.txt rules (the standard files websites use to tell crawlers what they may or may not access). One publishing executive, Lindsay Van Kirk of People Inc., described going from blocking roughly 2,100 user agents to over 30,000 after adopting a block-all-bots strategy — equivalent to tens of millions of scrape attempts per day. Media analyst Matthew Scott Goldstein estimates the broader scraper economy has grown into at least a $1 billion industry, with some executives suggesting it is actually multi-billion dollar in scale.

Analysis

Showing the shorter version.

Stealth Bot Traffic Is Up 597%. For Ad Revenue, That's Probably the Wrong Number to Panic About.

Cloudflare reports more than half of all web traffic is now bot-based. HUMAN Security, which sells bot mitigation, clocked stealth scraper traffic up 597% across 2025. Analyst Matthew Scott Goldstein puts the scraper economy at $1 billion. Those are big numbers. They are also, in part, a vendor sales deck.

The more useful question is what these scrapers actually touch. Stealth crawlers pull HTML to harvest text for AI training. They mostly don't render the JavaScript that fires an ad impression. That means they are stealing your content, not your ad revenue. DoubleVerify and Integral Ad Science (the two dominant impression-verification vendors) have not flagged a material spike in invalid-traffic rates tied to AI crawlers, and if scrapers were inflating measured impressions, those companies would be the loudest voices in the room. It sells product.

That said, there is a slower measurement problem that deserves real attention. If bot-inflated pageviews widen the gap between reported traffic and actual human reach, open-web CPMs can look stable while the real audience quietly shrinks. Buyers eventually reconcile those numbers. When they do, brand budgets shift toward authenticated inventory: direct deals, Reddit, the New York Times, CTV, places where a logged-in human is a logged-in human.

Lindsay Van Kirk of People Inc. shows what the operational reality looks like: her team went from blocking 2,100 user agents to more than 30,000. That is a standing ops function, not an IT ticket, and nobody budgeted for it. The risk in aggressive blocking is that you catch Googlebot or a partner's crawler in the net and torch real revenue to stop a scrape that never cost you a cent.

The other thing worth separating out is content theft. If OpenAI or Anthropic is training on your inventory, that is a licensing conversation, not a firewall problem. Most mid-tier publishers have no leverage to extract a licensing deal, so spending heavily on blocking infrastructure to solve a problem you can't monetize anyway is a questionable trade.

Before committing budget to mitigation: pull your own server logs and compare human-verified sessions against reported pageviews. If they have diverged more than a few points over the past year, you have a real measurement problem to fix before Q1 2026 forecasts go out. If they haven't, the headline number is mostly a vendor's fear.

Our call: By the end of Q1 2026 earnings calls (late April 2026), neither DoubleVerify nor IAS will report a material IVT spike attributable to AI scrapers. Scrapers don't render pages at scale; they don't need to. The impression-fraud panic and the content-theft problem are real, but they are separate, and most of the coverage is blending them.

Also covered this issue

Comments