blog.cloudflare.com web signal

Cloudflare will block AI training bots on ad pages Sept 15

4 sources tracking this story

TL;DR

  • Bot traffic has surpassed human traffic on the internet for the first time, lending structural weight to Cloudflare's timing beyond the policy itself.
  • The three-category taxonomy (Search, Agent, Training) is being positioned as a candidate industry standard via the Content Signals format, not only a Cloudflare dashboard feature.
  • Google, Apple, and Microsoft run mixed-use crawlers that bundle search indexing with training collection; publishers who block Training will also lose search visibility unless those companies separate the functions.

One thing worth watching in the AI-versus-publishers slow burn: Cloudflare has picked a date. On September 15, 2026, the company announced that new defaults will kick in for how AI crawlers are treated on sites behind its network. The old binary of block-AI-bots-or-don't is gone; in its place is a three-way split between Search, Training and Agent traffic.

The mechanics are the important part. Cloudflare defines Search as behavior that collects or indexes content to answer questions later, Agent as automated behavior acting in real time on a person's behalf, and Training as a crawler taking content to train or fine-tune a model. Under the new defaults, Training and Agent will be blocked on pages that display ads, while Search will remain allowed. The rules apply to all new domains onboarding to Cloudflare, new customers, and existing customers on the Free tier, with an opt-out window before September 15.

The framing from CEO Matthew Prince is the tell. "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge," he said, according to The Register. The pitch is that an ad is a signal the site owner meant a human to land there, so scraping that page to train a model or to feed an agentic answer is a different transaction from search indexing.

Two things ship alongside the deadline. BotBase, a database of known bots with their verified status and declared purpose, is a view Cloudflare says it hasn't shown dynamically on the dashboard before. And Pay Per Crawl is being rebranded to Pay Per Use, with launch partners Ceramic.ai paying publishers when content appears in search results and You.com paying when premium content is accessed by an agent.

The honest caveats. The reporting doesn't spell out how mixed-use crawlers from Google, Apple and Microsoft will actually behave when Training is blocked and their bots fall under the most restrictive rule, and it doesn't give you a rate card for what Pay Per Use will pay a small publisher per use. The Search / Training / Agent categories are also declarations by the AI companies themselves, so anyone tempted to label a training run as 'Search' isn't stopped by a fingerprint in the announced policy.

The opening this creates is for publishers who sat out earlier AI content deals because the terms were binary. If Pay Per Use pays anything meaningful and Ceramic.ai and You.com are the first two willing to sign, the interesting question is who is second.

What others are reporting

Coverage cluster as of 24h after publish

  1. TechCrunch Read →

    Frames the policy against the milestone that bot traffic now exceeds human traffic; names Ceramic.ai and You.com as Pay Per Use partners; reports Google's counter-argument about its separate opt-out crawler.

    Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge. — Matthew Prince, Cloudflare CEO
  2. Engadget Read →

    Sharpest framing of the Google asymmetry: Googlebot bundles search and training, so publishers cannot block training without accepting search-visibility risk — Cloudflare names the problem but cannot resolve it unilaterally.

    Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge. — Matthew Prince, Cloudflare CEO
  3. Help Net Security Read →

    Most technical depth of the three: names product execs Jin-He Lee and Bryan Becker, introduces the Content Signals format and transitive trust model as the standards layer beneath the September deadline.

    Content owners still want to be able to protect their content, and they should be compensated for the original content that they work hard to create, curate, and share.