WordPress robots.txt for AI crawlers: allow search, block training, and llms.txt
Published 7 October 20269 min readBy the SmoothSeen editorial team
The WordPress robots.txt is a virtual file generated by core that you can extend with the robots_txt filter, Yoast SEO or Rank Math. For AI, let the search crawlers in (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and, if you choose, block the training crawlers (GPTBot, ClaudeBot, Google-Extended). An llms.txt file is optional and has no proven effect.
Key points
- WordPress serves a virtual robots.txt built by do_robots and extensible through the robots_txt filter; a physical robots.txt file in the web root replaces it entirely.
- Since WordPress 5.3, the setting that discourages search engines adds a noindex robots meta tag and no longer writes Disallow into robots.txt.
- OAI-SearchBot, Claude-SearchBot, Claude-User and PerplexityBot decide whether ChatGPT, Claude and Perplexity can cite you; GPTBot, ClaudeBot, Google-Extended and Applebot-Extended control training.
- Yoast SEO (since 25.3, June 2025) and Rank Math can generate llms.txt, but Google says it does not need it and no AI search engine has documented reading it.
- A security plugin, the host or a firewall can answer 403 to a bot that robots.txt allows, and nothing in robots.txt shows it.
To check it on your own site: AI visibility audit
On this page
- What is in the robots.txt that WordPress generates?
- What does "Discourage search engines from indexing this site" do?
- Search versus training: what each crawler decides
- A commented robots.txt for WordPress
- How to edit robots.txt in WordPress
- Do you need an llms.txt on WordPress?
- When robots.txt says yes and the server says no
- How to check it
- What SmoothSeen does with this
- What to do this week
- Frequently asked questions
A robots.txt file is a plain-text file at the root of the domain that tells each crawler which paths it may request. It is not a lock: it only binds the crawlers that choose to honour it. On WordPress, core writes one for you, SEO plugins extend it, and the server, host or CDN can contradict it without you noticing. This guide explains what each layer controls and how to set the file up so AI search engines can cite you.
What is in the robots.txt that WordPress generates?
When there is no robots.txt file in the web root, WordPress answers that URL with the do_robots() function. In the current version, 7.1, this is the base output for a site that is visible to search engines1:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://www.example.com/wp-sitemap.xmlThe first two rules keep crawlers out of the dashboard while leaving admin-ajax.php open, because many themes call it from public pages. The Sitemap line comes from the core sitemap and only appears when the site is public1. If you run Yoast SEO or Rank Math, the sitemap listed is theirs.
Everything passes through the robots_txt filter, available since WordPress 3.0, which receives the text plus a second parameter saying whether the site is public2. It is the route plugins use to write into the file.
One detail causes endless confusion: if you upload a physical robots.txt to the web root, the server delivers it directly and WordPress never runs. Neither the filter nor any plugin can change it. Rank Math says so in its documentation: to edit the file from the plugin, delete the physical one first3.
What does "Discourage search engines from indexing this site" do?
Under Settings > Reading, the Search engine visibility section has a checkbox labelled Discourage search engines from indexing this site. Since WordPress 5.3 it no longer adds Disallow: / to robots.txt: it prints a <meta name='robots' content='noindex,nofollow' /> tag on every page instead4. The WordPress team's reasoning was that a site can still show up in results without being crawled, if other sites link to it4.
Two consequences follow:
- With the box ticked, your robots.txt still lets everyone in. The block lives in the HTML of each page.
- If you add rules through the
robots_txtfilter, check the second parameter: on a staging site with the box ticked you probably want to leave the file alone.
Search versus training: what each crawler decides
This is the part people mix up most. Some crawlers decide whether an assistant can cite you in its answers; others decide whether your content may be used to train models. Blocking the second group is a legitimate editorial choice and does not remove you from answers.
- Crawler
- OAI-SearchBot
- Company
- OpenAI
- What it is for, according to its owner
- ChatGPT search
- If you block it
- You are not shown in ChatGPT search answers5
- Crawler
- GPTBot
- Company
- OpenAI
- What it is for, according to its owner
- Model training
- If you block it
- Your content is not used for training; search is unaffected5
- Crawler
- Claude-SearchBot
- Company
- Anthropic
- What it is for, according to its owner
- Improving Claude's search results
- If you block it
- It may reduce your visibility in Claude6
- Crawler
- Claude-User
- Company
- Anthropic
- What it is for, according to its owner
- Fetching pages when a Claude user asks
- If you block it
- It may reduce your visibility in Claude6
- Crawler
- ClaudeBot
- Company
- Anthropic
- What it is for, according to its owner
- Model training
- If you block it
- Your future content is excluded from its training data6
- Crawler
- PerplexityBot
- Company
- Perplexity
- What it is for, according to its owner
- Surfacing and linking sites in its results; not used for training
- If you block it
- You do not appear in its results7
- Crawler
- Google-Extended
- Company
- What it is for, according to its owner
- robots.txt token for Gemini training and grounding in its apps
- If you block it
- No effect on your inclusion in Google Search8
- Crawler
- Applebot-Extended
- Company
- Apple
- What it is for, according to its owner
- Does not crawl: decides whether Apple may train on what Applebot already collected
- If you block it
- Your content is not used to train its models9
- Crawler
- CCBot
- Company
- Common Crawl
- What it is for, according to its owner
- Open archive of the web, also used by AI researchers
- If you block it
- You leave its archive10
Three caveats that change decisions. OpenAI says each setting is independent: you can allow OAI-SearchBot and block GPTBot5. ChatGPT-User acts on a user's request and OpenAI warns that robots.txt rules may not apply to it5; Perplexity says the same of Perplexity-User7. And AI Overviews and AI Mode are part of Google Search: they depend on Googlebot, not on Google-Extended11.
A commented robots.txt for WordPress
This example lets every crawler in, AI search crawlers included, except into the dashboard, and opts out of training:
# General rule. Googlebot, Bingbot and the AI search crawlers
# (OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot)
# fall here: they may crawl everything except the dashboard.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
# Model training: opted out (editorial decision).
# Blocking these does not remove you from ChatGPT, Claude,
# Perplexity or Google Search.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
Disallow: /
Sitemap: https://www.example.com/wp-sitemap.xmlWhy there is no separate group for the AI search crawlers: Google explains that each crawler obeys only the group with the most specific user agent that matches it12. Create a User-agent: OAI-SearchBot group and that bot stops reading the general group, so you would have to repeat every rule. As long as you want to treat them like Googlebot, the general group is enough.
How to edit robots.txt in WordPress
Pick one route. Combine a physical file, a plugin and a filter, and the result is easily not the one you think.
With Yoast SEO
Yoast's documentation points to Yoast SEO > Tools > File editor. The menu only appears when the file is writable and file editing has not been disabled on your install; otherwise Yoast suggests creating the file over FTP13.
With Rank Math
Switch on Advanced Mode and go to Rank Math SEO > General Settings > Edit robots.txt. Rank Math serves a virtual file, so you must first delete any physical robots.txt3.
With a PHP filter
If you do not use an SEO plugin, or prefer code, a must-use plugin on the robots_txt filter appends the training block to the core file. Save it as wp-content/mu-plugins/robots-ai.php:
<?php
/**
* Plugin Name: robots.txt for AI crawlers
* Description: Opts out of model training without touching AI search crawlers.
*/
if ( ! defined( 'ABSPATH' ) ) {
exit;
}
add_filter( 'robots_txt', function ( $output, $is_public ) {
// If the site is set to discourage search engines, leave it alone.
if ( ! $is_public ) {
return $output;
}
$training = array(
'GPTBot',
'ClaudeBot',
'Google-Extended',
'Applebot-Extended',
'CCBot',
'meta-externalagent',
);
$output .= "\n# Model training: opted out (editorial decision)\n";
foreach ( $training as $bot ) {
$output .= "User-agent: {$bot}\n";
}
$output .= "Disallow: /\n";
return $output;
}, 20, 2 );Several consecutive User-agent lines form a single group that shares the rules below them12. It has no effect if a physical robots.txt exists.
Do you need an llms.txt on WordPress?
llms.txt is a proposal by Jeremy Howard (September 2024) to publish at /llms.txt a Markdown summary of the site with links to its main pages, meant for language models to read14. It is not a standard and no AI search engine has documented using it. Google says you need no special text file for AI to appear in its AI features11, and in June 2026 it clarified that such files neither help nor hurt in Google Search, although you can keep them for other services15. The evidence on who actually requests the file is gathered in the llms.txt guide.
If you still want one, it is cheap:
- Yoast SEO generates it from version 25.3 (10 June 2025)16. According to its help pages: Yoast SEO > Settings > Site Features, AI tools section, LLMS.txt option17. You can let Yoast pick the content, giving priority to what you mark as cornerstone, or choose it yourself.
- Rank Math has an LLMS Txt module that you switch on under Rank Math SEO > Dashboard and configure under General Settings > Edit llms.txt18. Its documentation does not say which version introduced it, and the guide itself describes llms.txt as a proposal, not an official standard18.
Treat it as an optional extra: robots.txt and real crawler access come first.
When robots.txt says yes and the server says no
A robots.txt that allows OAI-SearchBot is worthless if the server answers it with a 403. On WordPress this happens more often than you would think:
- Security plugins with rules that block "suspicious" user agents or bot lists.
- The host's firewall, which sometimes blocks crawlers without telling you.
- The CDN. Cloudflare, for instance, has a setting to block AI bots; the guide to Cloudflare and AI bots walks through it.
A user-agent block on the server is as effective as one in robots.txt, but you will not see it by reading the file. The reverse also holds: blocking by user agent does not protect you from anyone who pretends to be another bot. Anthropic also advises against blocking by IP address, because it can stop its bot from reading your robots.txt6.
How to check it
- Open
https://www.example.com/robots.txtin a browser and read what is actually served. - Look at the robots.txt report in Search Console: it shows which file Google found, when it last crawled it and any errors, and lets you request a recrawl19.
- Request a page with a user agent that contains the bot's name and compare the status code with a browser's. It is an approximation: some firewalls also verify the IP.
curl -sI -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" https://www.example.com/
curl -sI https://www.example.com/What SmoothSeen does with this
SmoothSeen reads the robots.txt of the page you analyse and separates AI search crawlers from training crawlers: blocking the training ones is not counted as a failure. It also requests the page with the user agent of each AI crawler, a browser and Googlebot, to detect whether the server or firewall answers 403 to any of them, and checks whether an llms.txt exists. All of it feeds the AI visibility score.
What to do this week
Open your /robots.txt, make sure OAI-SearchBot, Claude-SearchBot and PerplexityBot are not blocked, and decide what to do about the training crawlers. Then test with curl that your server does not send them a 403. To have all of that access checked in one go, run an AI visibility audit.
Frequently asked questions
Where is the WordPress robots.txt if I cannot see it over FTP?
Nowhere on disk: WordPress builds it on the fly whenever someone requests /robots.txt and no file with that name sits in the web root. That is why it does not show up in FTP or in your host's file manager. To change it, use an SEO plugin or the robots_txt filter, or create a physical file, knowing that WordPress then stops having a say.
If I block GPTBot, do I disappear from ChatGPT?
No. GPTBot is the crawler OpenAI uses to train models, and its documentation says each setting is independent. What decides whether you appear in ChatGPT search answers is OAI-SearchBot. You can block GPTBot and allow OAI-SearchBot, a common combination for publishers who want to be cited without handing their content over for training.
Does Google-Extended affect AI Overviews or AI Mode?
No. Google-Extended is a robots.txt token that controls whether Google uses your content to train Gemini and for grounding in its apps, and Google states it does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are part of Search and rely on Googlebot; to limit them, Google points to nosnippet and noindex.
Do Yoast SEO or Rank Math create llms.txt automatically?
Only once you switch the feature on. In Yoast SEO it sits in the site features settings, under the AI tools heading, and the free version includes it. In Rank Math it is a module you enable from its dashboard. Either way the file is published at the root of the domain; remember that it is optional and Google does not use it.
Sources
- 1WordPress 7.1 source: functions.php (do_robots) and class-wp-sitemaps.php, WordPress (official wordpress-develop repository), accessed 7 October 2026.
- 2do_robots(), WordPress Developer Resources, accessed 7 October 2026.
- 3Using Rank Math's robots.txt generator, Rank Math, accessed 7 October 2026.
- 4Changes to prevent search engines indexing sites, Make WordPress Core, 2 September 2019.
- 5Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 6Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
- 7Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 8Google's common crawlers, Google Search Central, updated 14 July 2026.
- 9About Applebot, Apple, updated 4 September 2026.
- 10CCBot, Common Crawl, accessed 7 October 2026.
- 11AI features and your website, Google Search Central, updated 10 December 2025.
- 12How Google interprets the robots.txt specification, Google Search Central, updated 31 August 2026.
- 13How to edit robots.txt through Yoast SEO, Yoast, accessed 7 October 2026.
- 14The /llms.txt file, Jeremy Howard, 3 September 2024.
- 15Latest Google Search documentation updates, Google Search Central, updated 1 October 2026.
- 16Yoast SEO 25.3 changelog, Yoast, 10 June 2025.
- 17How to enable llms.txt with Yoast SEO, Yoast, accessed 7 October 2026.
- 18How to use llms.txt in Rank Math SEO, Rank Math, accessed 7 October 2026.
- 19robots.txt report, Search Console Help, accessed 7 October 2026.
How to cite this article
SmoothSeen. (2026, October 7). WordPress robots.txt for AI crawlers: allow search, block training, and llms.txt. https://smoothseen.com/en/blog/wordpress-robots-txt-llms-txt/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.