Cloudflare: block AI bots and AI training without vanishing from ChatGPT or Google
Published 7 October 202610 min readBy the SmoothSeen editorial team
To block AI bots in Cloudflare without losing visibility, open Security Settings, set Training to "Disallow AI Training", leave Search on "Allow" and decide on Agent separately. Do not pick "Block" for Training: since 15 September 2026 that option also shuts out Googlebot, Bingbot and Applebot. Blocking Search removes you from ChatGPT, Claude and Perplexity.
Key points
- Since 1 July 2026 Cloudflare sorts AI bots into three behaviours (Search, Agent and Training) and lets every plan, including Free, block each one separately.
- Since 15 September 2026, setting Training to "Block" also blocks Googlebot, Bingbot and Applebot, because they crawl for search and training at once. "Disallow AI Training" is the setting that stops training only.
- OAI-SearchBot, Claude-SearchBot and PerplexityBot are search crawlers. Block them and ChatGPT, Claude and Perplexity can no longer cite you. GPTBot, ClaudeBot and CCBot are training crawlers, and blocking them does not remove you from answers.
- robots.txt states a preference; Cloudflare's firewall returns a 403. A firewall block never shows up in robots.txt.
To check it on your own site: AI visibility audit
On this page
- What does Cloudflare control, and what does it not?
- Search, Agent and Training: the three behaviours
- What changed on 15 September 2026
- Decision table: what you want and which setting to choose
- AI Crawl Control: allow or block crawler by crawler
- A WAF rule that lets search through and stops training
- robots.txt, Content Signals and Bot Preference Sync
- How to check it
- What SmoothSeen does with this
- What to do this week
- Frequently asked questions
Cloudflare is a reverse proxy: every request to your site goes through its network before it reaches your server, so it can answer a crawler with a 403 without your server ever knowing. That makes it a handy way to keep your content out of model training, and a risky one if you tick the wrong box. This guide sets out what each setting does according to Cloudflare's documentation and official blog as of 7 October 2026, because the controls changed twice in three months.
What does Cloudflare control, and what does it not?
Cloudflare works at two separate layers, preference and enforcement, and they are easy to mix up:
- Layer
- robots.txt (managed by Cloudflare or your own)
- What it does
- Publishes a preference: "do not train on this"
- Who honours it
- Only crawlers that choose to comply
- Layer
- AI bot settings and AI Crawl Control
- What it does
- Block the request on Cloudflare's network with a 403
- Who honours it
- Every bot Cloudflare identifies
- Layer
- WAF custom rules
- What it does
- Block, challenge or skip according to the expression you write
- Who honours it
- Every request that matches
Cloudflare's own documentation says robots.txt compliance is voluntary and the file does not stop access at a technical level1. The firewall does. The consequence people forget: a firewall block does not appear in robots.txt. Your robots.txt can welcome OAI-SearchBot while a Cloudflare rule turns it away with a 403.
What Cloudflare does not control is whether an assistant cites you. That depends on the page itself. Cloudflare only decides who gets in. If your site does not sit behind Cloudflare, the same user-agent block can be done on the server, in Apache with .htaccess or in Nginx. The rest of the Cloudflare set-up (HTTPS, HSTS, headers and caching) is covered in the guide to Cloudflare security headers and caching.
Search, Agent and Training: the three behaviours
Since 1 July 2026, every Cloudflare customer, Free plan included, can manage three AI behaviours separately2:
- Search: crawlers that collect or index your content so they can answer questions about it later.
- Agent: automated activity acting in real time on a person's behalf, such as the bots that open a page when someone asks for it in a chat.
- Training: crawlers that take your content to train or fine-tune a model.
Each one can be set to "Block" on every page, "Block on pages with ads" (only where Cloudflare detects advertising) or "Allow"2. A single bot can have more than one behaviour3. This is how Cloudflare's reference classifies the most common crawlers, next to what each operator says about them:
- Crawler
- OAI-SearchBot
- Cloudflare category
- AI Search
- What it is for, according to its operator
- ChatGPT search: opt out and you are not shown in its answers4
- Crawler
- Claude-SearchBot
- Cloudflare category
- AI Search
- What it is for, according to its operator
- Claude's search index5
- Crawler
- PerplexityBot
- Cloudflare category
- AI Search
- What it is for, according to its operator
- Surfacing and linking sites in Perplexity results6
- Crawler
- ChatGPT-User, Claude-User, Perplexity-User
- Cloudflare category
- AI Assistant
- What it is for, according to its operator
- Fetching a page when a user asks for it
- Crawler
- GPTBot, ClaudeBot, meta-externalagent
- Cloudflare category
- AI Crawler
- What it is for, according to its operator
- Model training (Cloudflare puts CCBot and Bytespider in the same category)
- Crawler
- Googlebot, Bingbot
- Cloudflare category
- Search Engine
- What it is for, according to its operator
- Traditional search; Cloudflare treats them as mixed-use (search and training)
The categories in the second column come from Cloudflare's bot reference7. Two details matter. Blocking Agent stops ChatGPT, Claude and Perplexity from opening your page when a user asks for it mid-conversation, and OpenAI and Perplexity both warn that these user-triggered fetchers may not follow robots.txt46. And Google-Extended is missing from every blocking table because it is not a crawler: it is a robots.txt token with no user agent of its own8. You can only control it in robots.txt.
What changed on 15 September 2026
Some background: on 1 July 2025 Cloudflare announced it was switching to blocking AI crawlers by default unless they paid for content9. A year later it changed approach and split the three behaviours. The outstanding problem was a different one: Google, Microsoft and Apple each use a single crawler for search and for training. Until September, Cloudflare's Training block left them alone. Since 15 September 2026 it does not10:
- "Block" and "Block on pages with ads" for Training now apply to mixed-use crawlers, including Googlebot, Bingbot and Applebot. Choosing "Block" stops them entirely, search included.
- "Disallow AI Training" is new: it publishes a no-training preference in robots.txt, lets through the mixed-use crawlers Cloudflare labels "Accountable" (Googlebot, Bingbot and Applebot) and blocks every other training crawler, including the training-only crawlers run by Amazon, Anthropic, Meta and OpenAI.
- The old "Block AI bots" toggle is being retired in favour of the three controls, and managed robots.txt gives way to Bot Preference Sync, which writes into robots.txt whatever you configure in the dashboard11.
- New domains choose between two recommended presets: everything on "Allow" if the site does not run ads, or Training on "Disallow AI Training" and Agent blocked on ad pages if it does.
Cloudflare says customers on the old settings are migrated automatically and need do nothing10. Check what you ended up with anyway. One caveat for Bing: Microsoft does not yet read a no-training preference in robots.txt (it is targeting early 2027), so "Disallow AI Training" does not pass that preference on to Bing for now10.
A useful benchmark: according to Cloudflare, fewer than 1% of its sites block Search, while 17% use some mechanism to block training (September 2026)10.
Decision table: what you want and which setting to choose
- What you want
- Be cited by ChatGPT, Claude, Perplexity and Google, but keep your site out of training
- Search
- Allow
- Training
- Disallow AI Training
- Agent
- Allow
- What you want
- The same, and keep agents off your ad-funded pages
- Search
- Allow
- Training
- Disallow AI Training
- Agent
- Block on pages with ads
- What you want
- Maximum visibility, no restrictions
- Search
- Allow
- Training
- Allow
- Agent
- Allow
- What you want
- Disappear from every search engine, Google included
- Search
- Block
- Training
- Block
- Agent
- Block
Hardly anyone wants the last row, and it is exactly what you get by accident if you pick "Block" with AI in mind. The settings live in your zone's Security Settings10. If your site has one business goal (being found), the first row is the sensible one: training is held back and visibility stays intact.
AI Crawl Control: allow or block crawler by crawler
AI Crawl Control (formerly AI Audit) is Cloudflare's dashboard for seeing which AI crawlers visit your site and deciding on each one12. You will find it in the zone menu under AI Crawl Control, on the Crawlers tab, in the Action column: "Allow" or "Block". Three points from the documentation1213:
- On the Free plan it identifies crawlers by user agent; the finer detection based on Bot Management detection IDs is an Enterprise feature.
- Blocking creates or updates a single WAF custom rule named "AI Crawl Control". On paid plans you can choose whether it returns a 403 or a 402.
- If a crawler set to "Allow" is still blocked, look for an earlier WAF rule stopping it: the AI Crawl Control rule only lists blocked crawlers.
The third option, Pay per crawl, is still in private beta as of 7 October 202612.
A WAF rule that lets search through and stops training
If you need finer control than the three settings (for example, blocking training only under /blog/), use custom rules under Security rules > Create rule > Custom rules14. These expressions use fields that exist in Cloudflare's reference: cf.client.bot (verified bot) and cf.verified_bot_category (its category)3.
First rule, action Skip ("All remaining custom rules"), placed at the top:
(cf.client.bot and cf.verified_bot_category in {"Search Engine Crawler" "AI Search"})Second rule, action Block:
(cf.client.bot and cf.verified_bot_category eq "AI Crawler")
or (http.user_agent contains "GPTBot")
or (http.user_agent contains "ClaudeBot")
or (http.user_agent contains "CCBot")
or (http.user_agent contains "Bytespider")
or (http.user_agent contains "meta-externalagent")Why it is written this way:
- The "AI Crawler" category does not include Googlebot (that is "Search Engine Crawler") or OAI-SearchBot ("AI Search"), so this rule leaves search alone7.
- The user-agent lines catch crawlers that are not verified.
containsis case-sensitive and "ClaudeBot" is not part of "Claude-SearchBot", so search is not caught by mistake. - The original categories still work in WAF rules, according to Cloudflare, even though its newer taxonomy uses Search, Agent and Training3.
The Free plan allows five custom rules, and AI Crawl Control uses one of them once you block anything with it13. A warning: blocking by user agent does not stop anyone who fakes it. A scraper posing as Chrome walks straight past the second rule.
robots.txt, Content Signals and Bot Preference Sync
Content Signals are three preferences Cloudflare writes into robots.txt: search (building a search index and showing linked results), ai-input (feeding content into a model to answer, as in RAG) and ai-train (training or fine-tuning models)1. Cloudflare's managed robots.txt places this ahead of your own file:
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /Below it comes a Disallow: / for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent1. None of them is a search crawler. Note that it says nothing about ai-input, so real-time answers from an assistant are left without a stated preference.
Two warnings from the documentation. Search Console may flag those lines as "Syntax not understood", and Cloudflare says it has seen no effect on crawling or SEO1. And since September, managed robots.txt is being replaced by Bot Preference Sync, which is on by default for new customers and prepends its lines to your robots.txt without removing your own Disallow rules11. If you rely on fine-grained WAF rules, Cloudflare suggests switching the sync off and writing the file yourself, because it does not read custom rules11.
How to check it
- Look at the robots.txt that is actually served, not the one on your server: open
https://example.com/robots.txtand look for theBEGIN Cloudflareblock. - Open AI Crawl Control > Metrics: it shows response codes per crawler. If OAI-SearchBot is getting 403s, something is blocking it.
- Test with
curl, requesting the page with each bot's user agent:
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://example.com/
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://example.com/The curl test has a limit worth understanding: Cloudflare verifies bots by IP address, reverse DNS or cryptographic signature, not by user agent alone3. A request from your laptop carrying OAI-SearchBot's user agent is not a verified bot. If the block depends on the user agent (AI Crawl Control on the Free plan, or the second rule above), curl reproduces it. If it depends on the verified-bot category, you will only see it in Cloudflare's metrics.
What SmoothSeen does with this
SmoothSeen requests your page with each AI crawler's user agent and compares the result with what a browser and Googlebot receive. If the firewall returns a 403 to OAI-SearchBot or PerplexityBot while serving the page to everyone else, it flags it, even when robots.txt lets them in. It also reads the robots.txt that is actually served and keeps search crawlers apart from training crawlers; blocking GPTBot or ClaudeBot is not counted as a failure. For the reason explained in the previous section, a block that only applies to bots verified by IP may not be reproducible from outside.
What to do this week
Open your zone's Security Settings and note what Search, Training and Agent are set to. If Training says "Block", switch it to "Disallow AI Training" unless you also want to drop out of Google. Then check from the outside what each AI crawler receives with an AI visibility audit.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT?
No. According to OpenAI, GPTBot is the training crawler, while ChatGPT search relies on OAI-SearchBot, which is controlled separately. You can block GPTBot in robots.txt or in Cloudflare and still appear in ChatGPT's search answers, as long as OAI-SearchBot can get in and no firewall rule hands it a 403.
Does Cloudflare block AI bots by default?
It depends on when and how your domain was added. In July 2025 Cloudflare announced it would block AI crawlers by default. Since 15 September 2026, new domains choose between two presets: no blocks for sites that do not run ads, or "Disallow AI Training" plus agents blocked on ad pages for sites that do. Always check Security Settings.
What is the difference between "Block" and "Disallow AI Training"?
"Block" on Training shuts out every crawler that trains, including those that also power search, such as Googlebot, Bingbot and Applebot. "Disallow AI Training" writes the no-training preference into robots.txt, lets those three through for search and blocks every other training crawler, such as GPTBot or ClaudeBot.
If robots.txt allows a bot, can it still be blocked?
Yes. robots.txt and the firewall are separate layers. The crawler reads that it may enter, requests the page and Cloudflare answers with a 403 because of an AI bot setting, an AI Crawl Control rule or a WAF custom rule. From the outside you only notice by requesting the page as that bot or by checking AI Crawl Control's metrics.
Sources
- 1robots.txt setting, Cloudflare Docs, updated 3 August 2026.
- 2Block AI Bots, Cloudflare Docs, updated 1 July 2026.
- 3Verified bots, Cloudflare Docs, updated 1 July 2026.
- 4Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 5Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, accessed 7 October 2026.
- 6Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 7Bot reference, Cloudflare AI Crawl Control Docs, updated 23 April 2026.
- 8Google's common crawlers, Google Search Central, updated 14 July 2026.
- 9Content Independence Day: no AI crawl without compensation!, Matthew Prince, The Cloudflare Blog, 1 July 2025.
- 10Have it both ways: stay discoverable in search while disallowing AI training, Bryan Becker, The Cloudflare Blog, 15 September 2026.
- 11Say it once: introducing Bot Preference Sync, Jin-Hee Lee, The Cloudflare Blog, 21 August 2026.
- 12Manage AI crawlers, Cloudflare AI Crawl Control Docs, updated 28 July 2026.
- 13AI Crawl Control with Cloudflare WAF, Cloudflare Docs, updated 3 August 2026; plan limits in Custom rules, accessed 7 October 2026.
- 14Create a custom rule in the dashboard, Cloudflare Docs, updated 3 August 2026.
How to cite this article
SmoothSeen. (2026, October 7). Cloudflare: block AI bots and AI training without vanishing from ChatGPT or Google. https://smoothseen.com/en/blog/cloudflare-block-ai-bots/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.