// AI NEWS · PLATFORM

Cloudflare Lets Sites Block AI Training While Keeping Search Crawlers Allowed — But Googlebot's Dual Role Is the Catch

New per-purpose controls split AI bots into Search, Agent, and Training so a site can stay discoverable while refusing training crawlers. Because Googlebot does both jobs with one bot, blocking training can block Google Search too.

2026-09-16 · by Moe Ameen

What happened

Cloudflare has replaced its blunt "Block AI bots" toggle with per-purpose controls that let a site owner allow, or block, AI traffic by what the bot is actually doing. Announced on July 1, 2026, as a follow-up to its "Content Independence Day," the system — available to all customers, including the free plan — sorts crawlers into three categories a site can manage independently: Search, Agent, and Training. The headline capability is that a site can keep search bots allowed, so it stays discoverable, while blocking crawlers that only exist to train AI models.

Cloudflare defines the three categories by purpose. Search is any bot that indexes your content to answer questions about it later, with the expectation that you get referral traffic or comparable value in return. Agent is automated behavior acting in real time on a person's behalf — chat fetch bots like ChatGPT-User, and browser-use agents such as an AI driving Chrome. Training is a crawler that absorbs your content into a model's training data. The point of splitting them is that "block AI" was too coarse: many site owners want to stay in search results and answer engines while refusing to be free training fuel.

The catch is mixed-purpose crawlers, and Googlebot is the one that matters. Cloudflare classifies Googlebot (and Bingbot) as "Search + Training," because Google crawls for search indexing and AI purposes with a single bot rather than a separate, blockable training agent. When a bot serves multiple purposes, the most restrictive rule applies — so a site that sets AI Training to "block" can also block Googlebot. In the weeks before the change, site owners reported exactly this: enabling the training block returned HTTP 403 errors to Googlebot and Bingbot on their sitemaps, and Google's John Mueller publicly asked for details. In practice you cannot cleanly refuse Google's AI training while keeping its search crawl, because Google does not offer that separation.

Cloudflare also set new defaults taking effect September 15, 2026. For new domains, new sites created by existing customers, and all existing free-tier customers, bots classified as Training or Agent are blocked on pages that display ads, while Search stays allowed; mixed Search + Training crawlers get blocked by any configuration meant to block AI training, including the legacy "Block AI bots" option. Paying customers whose sites are already configured keep their current settings until they change them, and everyone can opt out of the new defaults through Cloudflare's security settings before the deadline. Confirm the exact category behavior and current defaults in Cloudflare's own documentation before changing anything, as the rules are still being refined.

Why it matters for creators

  • It turns a blunt on/off switch into a real choice: stay discoverable in search and answer engines while refusing training-only crawlers. That is the trade-off most creators actually want to make about their own site.
  • The Googlebot caveat is the trap. Because Google crawls for search and AI with one bot, a well-intentioned "block AI training" setting can also drop you out of Google Search — the opposite of what a creator wants.
  • Defaults are changing under people who never touch a setting. New sites and every free-tier site get the new behavior on September 15, 2026, so a hands-off site owner can have their crawl posture changed for them.
  • This is a control on ONE domain you own — your website. It governs what gets crawled there; it does nothing for the reach you build on platforms and in inboxes, where most creators actually grow.
  • The durable signal is that discoverability now splits between classic search and AI answer engines, and both are things you feed with original, well-structured content — not something a single crawler toggle can secure for you.

How to act on this with Kompozy

A Cloudflare toggle is defense on a single door: it decides what crawls your website. Useful — but it does nothing to make you discoverable in the places that actually drive growth, and getting it wrong can quietly drop you out of Google. The offense is producing enough original, well-structured content that you get found and cited everywhere — across search, AI answers, social feeds, and inboxes — so no one domain's crawl policy decides your reach. That production is what [Kompozy](/) exists to do. From one genuine idea it generates a depth-carrying [Blog Article](/glossary/output-buckets) and an [Email Newsletter](/glossary/output-buckets) — the structured, first-hand material that grounds both search rankings and AI answers — alongside brand-exact [Carousel Posts](/glossary/hyperframes), Quote Graphics, Photo Posts, and captioned shorts, all held to one voice by the [Persona Brief](/glossary/persona-brief).

The distribution is the point. Where a Cloudflare rule gates a crawler on your site, [Autopilot](/glossary/autopilot) does the opposite — it schedules and fans your set across the eight social platforms plus blog and email, behind a per-post review gate, so a single idea is indexed and surfaced in many owned and social places at once. Your newsletter list and blog are yours regardless of any crawler policy; your social presence keeps compounding whether or not a training bot ever reads your homepage. Set your Cloudflare categories however you like — allow search, block training, protect your ad pages — and let Kompozy make you big enough across nine destinations that the setting on one door stops being the thing your growth depends on.

Quick takeaways

  • Cloudflare replaced its "Block AI bots" toggle with per-purpose controls — Search, Agent, and Training — that all customers, including free-plan users, can allow or block independently. Announced July 1, 2026.
  • A site can now keep search crawlers allowed while blocking training-only and agent bots, so it stays discoverable without becoming free training data.
  • Googlebot and Bingbot are classified "Search + Training," and the most-restrictive rule wins — so blocking AI training can also block Google Search. Owners reported 403s to Googlebot on their sitemaps before the change.
  • From September 15, 2026, new domains, new sites by existing customers, and all free-tier customers get new defaults: Training and Agent blocked on ad pages, Search allowed. Paying, already-configured sites keep their settings until changed; opt-out is available.
  • The toggle governs one owned domain. Kompozy turns one idea into a blog, newsletter, carousels, and shorts and fans them across nine destinations, so discoverability is not hostage to a single crawler rule.

Frequently asked questions

Can I block AI training on my site but still let Google index it?

Cloudflare's new controls let you allow Search bots while blocking Training bots by purpose. The complication is Googlebot: Cloudflare classifies it as "Search + Training" because Google uses one bot for both, and the most-restrictive rule applies — so setting AI Training to "block" can also block Googlebot and drop you out of Google Search. For a training-only crawler the separation works; for Google it does not, because Google does not offer a separate, blockable training agent. Confirm current behavior in Cloudflare's docs.

What changes on September 15, 2026?

Cloudflare sets new defaults for new domains, new sites created by existing customers, and all existing free-tier customers: bots classified as Training or Agent are blocked on pages that display ads, while Search remains allowed. Mixed Search + Training crawlers are blocked by any configuration meant to block AI training, including the legacy "Block AI bots" option. Paying customers whose sites are already configured keep their current settings until they change them, and anyone can opt out of the new defaults in security settings beforehand.

What are the three Cloudflare AI bot categories?

Search is a bot that indexes your content to answer questions about it later, with an expectation of referral traffic or comparable value. Agent is automated behavior acting in real time on a person's behalf, such as ChatGPT-User or a browser-use agent driving Chrome. Training is a crawler that absorbs your content into a model's training data. Website owners can allow or block each category independently, including on the free plan.

Related news

← All AI news · Get started →