Cloudflare Just Declared War on Free-Riding AI Crawlers
Starting September 15, mixed-use crawlers—bots that harvest your content for both search indexing and AI training without distinguishing between the...

Cloudflare just drew a line in the sand that every AI company building on other people's content needs to pay attention to.
On September 15, 2026, Cloudflare will begin automatically blocking "mixed-use" crawlers from ad-supported websites hosted on its platform. These are bots that don't clearly separate why they're scraping a page—for search indexing, for AI model training, or for powering an AI agent. The company's position is blunt: if you can't tell publishers what you're using their content for, you don't get access.
The Crawler Economy Gets Complicated
For years, AI companies have operated under an implicit assumption: if a page is publicly accessible, scraping it is fine. The same crawler would index content for search, dump it into a training set, and later use it to power real-time agentic queries—all without asking permission or compensating the source. Publishers, caught in the middle, had almost no control beyond adding a robots.txt line that most crawlers ignored or interpreted loosely.
Cloudflare's move breaks that status quo. The company's network sits in front of roughly 20% of the web, which means its default settings carry enormous weight. By splitting crawler categories into three distinct buckets—Search, Training, and Agent—Cloudflare is forcing AI companies to be explicit about intent. If you want to train a model on a publisher's content, you'll need a Training-classified crawler. If you want to power an agent that reads pages in real time, that needs its own declared purpose. A single bot doing all three? Blocked.
Publishers Finally Get a Say
The policy shift is also a direct shot at the ad-supported publishing model that AI companies have been quietly undermining. When a crawler harvests a page to train a model or power a competitor's AI product, the publisher gets nothing—no referral traffic, no licensing fee, no acknowledgment. But the page carries ads. The crawler consumes bandwidth. The publisher foots the bill while someone else extracts the value.
Cloudflare's new controls give publishers the ability to allow Search bots (which drive traffic) while blocking Training and Agent bots by default. New domains onboarding to Cloudflare will have Training and Agent blocked automatically on pages that display ads. Existing customers can opt in to the same settings with a single toggle.
The Industry Response Will Be Revealing
How AI companies react will tell us a lot about how serious they are about content relationships. Firms like OpenAI, Anthropic, and Google have various licensing deals with major publishers, but the long tail of web content—the blog posts, independent news sites, and reference pages that actually train and power many AI systems—has largely been taken without negotiation.
The pragmatic move is to comply: split crawlers, negotiate licensing frameworks, and give publishers real opt-in mechanisms. The defensive move is to route around Cloudflare's network entirely, which is possible but carries its own infrastructure costs and reputational exposure.
This Is Infrastructure Acting Like a Gatekeeper
What's notable here is that Cloudflare isn't making an ethical argument—it's making a technical one. The company isn't saying AI training is wrong. It's saying: if you want to use our network to access our customers' content, you have to be transparent about what you're doing. That's a contractual position, not a philosophical one, and it's much harder to argue against.
Whether other infrastructure providers follow will determine whether this becomes an industry norm or a single-company stunt. Cloudflare has the position to make it matter. The next 72 days will show if the AI industry treats content creators as partners or just free real estate.


