AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Published 7 October 202610 min readBy the SmoothSeen editorial team
You can block GPTBot in robots.txt with User-agent: GPTBot and Disallow: /, and all that does is stop OpenAI using your content to train its models. It does not remove you from ChatGPT search, which depends on a different crawler, OAI-SearchBot. The major AI companies now split their crawlers by purpose, so you can refuse training and still be cited.
Key points
- Blocking GPTBot in robots.txt only stops OpenAI training on your content; ChatGPT search depends on OAI-SearchBot.
- The major AI companies split their crawlers into three uses, search, user-requested visits and training, and each one is blocked separately.
- Google-Extended does not affect Google Search, but it also decides whether the Gemini app can ground its answers in your site.
- robots.txt is a request, not a lock, and agents acting on a user's behalf warn that they may not follow it.
- A firewall, CDN or WAF can return a 403 to these crawlers even when your robots.txt lets them in; you can test it with curl and their user agent.
To check it on your own site: AI visibility audit
On this page
- Which AI crawlers exist and what does each one do?
- Why are search, user requests and training different?
- A robots.txt that lets you be cited without being trained on
- robots.txt is a request, not a lock
- The block that is not in robots.txt: firewalls, CDNs and WAFs
- How to check with curl what each crawler sees
- What SmoothSeen checks here
- Frequently asked questions
- What to do next
The distinction matters because the opposite was repeated for years: that blocking "the AI bots" wiped you from the assistants. This guide sits within AI search optimization, where crawler access is the first block to check. Below is the full list, checked against each company's documentation on 7 October 2026, an example robots.txt and a way to see whether your server lets in what your robots.txt promises.
Which AI crawlers exist and what does each one do?
An AI crawler is a program that requests web pages on behalf of an artificial intelligence company and identifies itself with its own name in the User-agent header. That name is what you write in robots.txt. Each company publishes its own and explains what it uses them for:
- Company
- OpenAI
- Crawler
- OAI-SearchBot
- Use
- Search: showing sites in ChatGPT answers
- Follows robots.txt?
- Yes
- Company
- OpenAI
- Crawler
- ChatGPT-User
- Use
- User-requested, not automatic crawling
- Follows robots.txt?
- OpenAI warns that its rules may not apply
- Company
- OpenAI
- Crawler
- GPTBot
- Use
- Training its models
- Follows robots.txt?
- Yes
- Company
- OpenAI
- Crawler
- OAI-AdsBot
- Use
- Checking pages submitted as ads in ChatGPT; no training
- Follows robots.txt?
- Not stated in the documentation
- Company
- Anthropic
- Crawler
- Claude-SearchBot
- Use
- Search: indexing for Claude's results
- Follows robots.txt?
- Yes
- Company
- Anthropic
- Crawler
- Claude-User
- Use
- User-requested, when Claude looks up a page
- Follows robots.txt?
- Yes, according to Anthropic
- Company
- Anthropic
- Crawler
- ClaudeBot
- Use
- Training
- Follows robots.txt?
- Yes
- Company
- Perplexity
- Crawler
- PerplexityBot
- Use
- Search: linking sites in its results; no training
- Follows robots.txt?
- Yes
- Company
- Perplexity
- Crawler
- Perplexity-User
- Use
- User-requested
- Follows robots.txt?
- Generally ignores it
- Company
- Crawler
- Googlebot
- Use
- Google Search, including AI Overviews and AI Mode
- Follows robots.txt?
- Yes
- Company
- Crawler
- Google-Extended
- Use
- Not a crawler: a robots.txt token for Gemini training and grounding
- Follows robots.txt?
- Yes
- Company
- Crawler
- Google-Agent and others
- Use
- User-requested
- Follows robots.txt?
- Generally ignore it
- Company
- Apple
- Crawler
- Applebot
- Use
- Search in Spotlight, Siri and Safari; its data may train Apple models
- Follows robots.txt?
- Yes
- Company
- Apple
- Crawler
- Applebot-Extended
- Use
- Does not crawl: decides whether Applebot's data trains Apple models
- Follows robots.txt?
- Yes
- Company
- Meta
- Crawler
- Meta-WebIndexer
- Use
- Improving Meta AI search results
- Follows robots.txt?
- Not stated in the documentation
- Company
- Meta
- Crawler
- meta-externalagent
- Use
- Training models and improving products
- Follows robots.txt?
- Yes
- Company
- Meta
- Crawler
- Meta-ExternalFetcher
- Use
- User-requested
- Follows robots.txt?
- May bypass it
- Company
- Amazon
- Crawler
- Amzn-SearchBot
- Use
- Search in Amazon products; no training
- Follows robots.txt?
- Yes
- Company
- Amazon
- Crawler
- Amazonbot
- Use
- Improving products; may train Amazon models
- Follows robots.txt?
- Yes
- Company
- Amazon
- Crawler
- Amzn-User
- Use
- User-requested, such as Alexa queries
- Follows robots.txt?
- May not follow all of its rules
- Company
- Common Crawl
- Crawler
- CCBot
- Use
- An open archive of the web that anyone can download
- Follows robots.txt?
- Yes, including
Crawl-delay
The sources for each row are the documentation from OpenAI1, Anthropic2, Perplexity3, Google45, Apple6, Meta7, Amazon8 and Common Crawl9. Names change and new crawlers appear, so it is worth reviewing the list every few months.
Why are search, user requests and training different?
Because each use has a different consequence for you, and the companies separate them precisely so that you can decide one by one.
Search. These crawlers read your site so the assistant can cite it in its answers. OpenAI is blunt about it: sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links1. Anthropic warns that blocking Claude-SearchBot may reduce your visibility in its results2. If you want to be cited, these need to get in. How each assistant works with them is covered in the guides to ChatGPT SEO and Perplexity SEO.
User requests. These are visits a person triggers from the assistant, for example by pasting a link and asking for a summary. They do not crawl on their own, which is why several companies warn that they may not follow robots.txt: OpenAI with ChatGPT-User, Perplexity with Perplexity-User, Google with its agents and Meta with Meta-ExternalFetcher1357.
Training. GPTBot, ClaudeBot, meta-externalagent and Amazonbot collect content that may be used to train models. Blocking them is a legitimate editorial decision: OpenAI makes clear that each setting is independent and that you can allow OAI-SearchBot while disallowing GPTBot1.
Two cases do not quite fit. Google-Extended has no user agent of its own: Google crawls with its usual agents and uses this token to know whether it may train Gemini on your content and use it to ground answers in the Gemini app and in Grounding with Google Search on Vertex AI. Google adds that it does not affect your inclusion in Google Search and is not a ranking signal4. Applebot-Extended does not crawl either: blocking it removes your content from Apple's training, but your pages can still appear in Spotlight, Siri and Safari6.
A robots.txt that lets you be cited without being trained on
This example lets each assistant's search crawler in and blocks training. It uses example.com, the domain reserved for examples; swap the private paths for your own:
# 1. AI search: they read you in order to cite you in their answers.
# They have their own group, so they do not inherit the "*" rules:
# the private paths are repeated here.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Amzn-SearchBot
User-agent: Meta-WebIndexer
Disallow: /basket/
Disallow: /my-account/
# 2. Model training and open archives: blocked.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /
# 3. Google-Extended blocks Gemini training, but also grounding
# in the Gemini app. Uncomment only if you accept that cost.
# User-agent: Google-Extended
# Disallow: /
# 4. Everyone else, including Googlebot: AI Overviews and AI Mode
# depend on it, not on a separate AI crawler.
User-agent: *
Disallow: /basket/
Disallow: /my-account/
Sitemap: https://www.example.com/sitemap.xmlThree decisions explain its shape. First, several consecutive User-agent lines share the rules that follow them, and under the robots.txt standard (RFC 9309) a crawler that finds a group naming it obeys that group and only falls back to * when no group names it10. That is why group 1 repeats the private paths. Second, the user-requested agents (ChatGPT-User, Perplexity-User) are left out because their owners do not commit to following the file; naming them gives a false sense of control. Third, Google-Extended is commented out because blocking it can reduce what Gemini cites from you.
One detail worth knowing if you use Cloudflare: its managed robots.txt adds a block on Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent in front of your own file11. That includes Google-Extended, with the Gemini cost you have just seen.
AI Overviews and AI Mode have no crawler or token of their own: they run on Googlebot and the usual snippet controls, as explained in the guide to Google AI Overviews.
robots.txt is a request, not a lock
Google says so in its introduction to the file: robots.txt instructions cannot enforce crawler behaviour, and it is up to each crawler to obey them12. RFC 9309 puts it another way: its rules are not a form of access authorisation10.
That has three practical consequences:
- Reputable crawlers respect it; others have no reason to. OpenAI, Anthropic, Perplexity, Google, Apple, Amazon and Common Crawl document that they follow its rules when crawling automatically. A crawler with no documentation, or one posing as a browser, owes you nothing.
- User-requested visits may skip it. OpenAI, Perplexity, Google and Meta say so about their own agents. If you need content that nobody reads, put it behind a password.
- robots.txt does not take a page out of an index. Google uses it to manage crawling, not to keep pages out of its search engine; that is what
noindexis for12. The technical SEO guide explains the difference in more detail.
It also takes time: OpenAI says its systems can need around 24 hours to reflect a robots.txt change in search1.
The block that is not in robots.txt: firewalls, CDNs and WAFs
Your robots.txt says what you declare; your server may do something else. Many hosts, CDNs and web application firewalls (WAFs) ship rules that answer certain user agents with a 403 Forbidden. If those rules include PerplexityBot or OAI-SearchBot, the crawler never even reads your robots.txt, and you see it in no dashboard.
The most widespread case is Cloudflare. On 1 July 2025 it announced it was changing its default to block AI crawlers unless they pay content creators13. Since 1 July 2026, all its customers can decide separately on three behaviours: search, agents acting on a person's behalf, and training. Since 15 September 2026, new domains start with training and agents blocked on pages that display ads, and with search allowed14. Its documentation adds another nuance: a crawler used for both search and training is blocked by any configuration that blocks training14.
If your site sits behind Cloudflare or another CDN, check those settings before trusting your robots.txt. OpenAI also asks that your host accept the IP addresses it publishes for each crawler, in files such as openai.com/searchbot.json1.
How to check with curl what each crawler sees
With curl you can request your page posing as a crawler and compare the response with a browser's. Replace example.com with your domain:
# 1. Control: an ordinary browser. Should return 200.
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/128.0 Safari/537.36" https://www.example.com/
# 2. The search crawlers, using the token each company publishes.
for ua in OAI-SearchBot Claude-SearchBot PerplexityBot; do
printf "%s: " "$ua"
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (compatible; $ua/1.0)" https://www.example.com/
done
# 3. What does your robots.txt say? Read what is actually served,
# including anything your CDN adds in front.
curl -s https://www.example.com/robots.txtHow to read the result:
- Browser
- 200
- Crawler
- 200
- What it means
- No rule filters on that name. It does not guarantee the real crawler gets in if the firewall also checks the IP
- Browser
- 200
- Crawler
- 403, 401, 406 or 451
- What it means
- A rule rejects that user agent. Look for it in your CDN, your WAF or your hosting configuration
- Browser
- 403 or 503
- Crawler
- Any
- What it means
- The server also rejects the browser from your network: the test is inconclusive
- Browser
- 200
- Crawler
- 429 or 5xx
- What it means
- Rate limit or server error: repeat later before drawing conclusions
A 403 for the crawler alongside a 200 for the browser is the clearest sign of a block outside robots.txt. To confirm that the real crawler gets in, search your server's access logs for its name and match the IP addresses against the official lists.
What SmoothSeen checks here
SmoothSeen reads your robots.txt and shows you, crawler by crawler, whether it lets in the search crawlers, the user-requested agents and the training crawlers, each group with its consequence. Blocking training is not presented as a failure, because it is not one. It also requests the page with each crawler's user agent, compares it with the response to a browser and warns you if the server rejects any of them even though robots.txt allows it. You will find it in the AI search half of the analysis.
What nobody can claim is that "all" AI bots are allowed or blocked: there is no complete register, and new crawlers appear every few months. SmoothSeen checks a closed list of documented crawlers and tells you which ones.
Frequently asked questions
Does blocking GPTBot remove me from ChatGPT?
No. GPTBot is OpenAI's training crawler: disallowing it indicates that your content should not be used to train its models. What decides whether you appear in ChatGPT search answers is OAI-SearchBot, which is configured separately. OpenAI uses exactly this as an example in its documentation: you can allow OAI-SearchBot to appear in search while disallowing GPTBot at the same time.
Does blocking Google-Extended affect AI Overviews?
No, according to Google. Google-Extended does not affect your inclusion in Google Search and is not a ranking signal, and AI Overviews and AI Mode are part of Search, which Googlebot crawls. What it does control, besides Gemini training, is whether the Gemini app and Grounding with Google Search on Vertex AI can use your content to ground their answers. That is the real cost of blocking it.
How do I block all AI bots at once?
No single line does it. robots.txt works by name, every company uses its own names, and they change. You can block the list in this guide, but crawlers that do not identify themselves or ignore the file will still get in. To actually stop them you need a rule in your CDN or firewall, or to put the content behind a password.
Why does my robots.txt allow PerplexityBot and it still does not cite me?
There may be a block outside the file. If a firewall, CDN or your host answers PerplexityBot's user agent with a 403, the crawler never reads the page even though robots.txt allows it. Test it with curl by comparing the response to a browser with the response to PerplexityBot. If both return 200, the reason lies elsewhere: the content, indexing or the competition.
What to do next
Open your robots.txt as it is actually served, check that no group blocks OAI-SearchBot, Claude-SearchBot or PerplexityBot, and decide separately what to do about training. Then run the curl test to rule out a server-side block. The other blocks that decide whether you get cited are in the guide to AI search optimization.
Sources
- 1Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 2Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic (Claude Help Center), updated 7 April 2026.
- 3Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 4Google's common crawlers, Google Crawling Infrastructure, updated 14 July 2026.
- 5Google User-Triggered Fetchers, Google Crawling Infrastructure, updated 19 August 2026.
- 6About Applebot, Apple, updated 4 September 2026.
- 7Meta Web Crawlers, Meta for Developers, accessed 7 October 2026.
- 8About AmazonBot, Amazon Developer, accessed 7 October 2026.
- 9CCBot, Common Crawl, accessed 7 October 2026.
- 10RFC 9309: Robots Exclusion Protocol, IETF, published September 2022.
- 11robots.txt setting, Cloudflare Docs, updated 3 August 2026.
- 12Introduction to robots.txt, Google Search Central, updated 10 December 2025.
- 13Content Independence Day: no AI crawl without compensation!, Matthew Prince, Cloudflare Blog, 1 July 2025.
- 14Block AI Bots, Cloudflare Docs, updated 1 July 2026.
How to cite this article
SmoothSeen. (2026, October 7). AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers. https://smoothseen.com/en/blog/ai-crawlers-robots-txt/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
ChatGPT SEO: how to rank in ChatGPT answers and recommendations
How to rank in ChatGPT, according to OpenAI: which crawler to allow, how it picks sources and products, and what makes it recommend a business.