Skip to content

WordPress robots.txt for AI crawlers: allow search, block training, and llms.txt

Published 7 October 20269 min readBy the SmoothSeen editorial team

The WordPress robots.txt is a virtual file generated by core that you can extend with the robots_txt filter, Yoast SEO or Rank Math. For AI, let the search crawlers in (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and, if you choose, block the training crawlers (GPTBot, ClaudeBot, Google-Extended). An llms.txt file is optional and has no proven effect.

Key points

  • WordPress serves a virtual robots.txt built by do_robots and extensible through the robots_txt filter; a physical robots.txt file in the web root replaces it entirely.
  • Since WordPress 5.3, the setting that discourages search engines adds a noindex robots meta tag and no longer writes Disallow into robots.txt.
  • OAI-SearchBot, Claude-SearchBot, Claude-User and PerplexityBot decide whether ChatGPT, Claude and Perplexity can cite you; GPTBot, ClaudeBot, Google-Extended and Applebot-Extended control training.
  • Yoast SEO (since 25.3, June 2025) and Rank Math can generate llms.txt, but Google says it does not need it and no AI search engine has documented reading it.
  • A security plugin, the host or a firewall can answer 403 to a bot that robots.txt allows, and nothing in robots.txt shows it.

To check it on your own site: AI visibility audit

On this page

A robots.txt file is a plain-text file at the root of the domain that tells each crawler which paths it may request. It is not a lock: it only binds the crawlers that choose to honour it. On WordPress, core writes one for you, SEO plugins extend it, and the server, host or CDN can contradict it without you noticing. This guide explains what each layer controls and how to set the file up so AI search engines can cite you.

What is in the robots.txt that WordPress generates?

When there is no robots.txt file in the web root, WordPress answers that URL with the do_robots() function. In the current version, 7.1, this is the base output for a site that is visible to search engines1:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.example.com/wp-sitemap.xml

The first two rules keep crawlers out of the dashboard while leaving admin-ajax.php open, because many themes call it from public pages. The Sitemap line comes from the core sitemap and only appears when the site is public1. If you run Yoast SEO or Rank Math, the sitemap listed is theirs.

Everything passes through the robots_txt filter, available since WordPress 3.0, which receives the text plus a second parameter saying whether the site is public2. It is the route plugins use to write into the file.

One detail causes endless confusion: if you upload a physical robots.txt to the web root, the server delivers it directly and WordPress never runs. Neither the filter nor any plugin can change it. Rank Math says so in its documentation: to edit the file from the plugin, delete the physical one first3.

What does "Discourage search engines from indexing this site" do?

Under Settings > Reading, the Search engine visibility section has a checkbox labelled Discourage search engines from indexing this site. Since WordPress 5.3 it no longer adds Disallow: / to robots.txt: it prints a <meta name='robots' content='noindex,nofollow' /> tag on every page instead4. The WordPress team's reasoning was that a site can still show up in results without being crawled, if other sites link to it4.

Two consequences follow:

  • With the box ticked, your robots.txt still lets everyone in. The block lives in the HTML of each page.
  • If you add rules through the robots_txt filter, check the second parameter: on a staging site with the box ticked you probably want to leave the file alone.

Search versus training: what each crawler decides

This is the part people mix up most. Some crawlers decide whether an assistant can cite you in its answers; others decide whether your content may be used to train models. Blocking the second group is a legitimate editorial choice and does not remove you from answers.

Crawler
OAI-SearchBot
Company
OpenAI
What it is for, according to its owner
ChatGPT search
If you block it
You are not shown in ChatGPT search answers5
Crawler
GPTBot
Company
OpenAI
What it is for, according to its owner
Model training
If you block it
Your content is not used for training; search is unaffected5
Crawler
Claude-SearchBot
Company
Anthropic
What it is for, according to its owner
Improving Claude's search results
If you block it
It may reduce your visibility in Claude6
Crawler
Claude-User
Company
Anthropic
What it is for, according to its owner
Fetching pages when a Claude user asks
If you block it
It may reduce your visibility in Claude6
Crawler
ClaudeBot
Company
Anthropic
What it is for, according to its owner
Model training
If you block it
Your future content is excluded from its training data6
Crawler
PerplexityBot
Company
Perplexity
What it is for, according to its owner
Surfacing and linking sites in its results; not used for training
If you block it
You do not appear in its results7
Crawler
Google-Extended
Company
Google
What it is for, according to its owner
robots.txt token for Gemini training and grounding in its apps
If you block it
No effect on your inclusion in Google Search8
Crawler
Applebot-Extended
Company
Apple
What it is for, according to its owner
Does not crawl: decides whether Apple may train on what Applebot already collected
If you block it
Your content is not used to train its models9
Crawler
CCBot
Company
Common Crawl
What it is for, according to its owner
Open archive of the web, also used by AI researchers
If you block it
You leave its archive10

Three caveats that change decisions. OpenAI says each setting is independent: you can allow OAI-SearchBot and block GPTBot5. ChatGPT-User acts on a user's request and OpenAI warns that robots.txt rules may not apply to it5; Perplexity says the same of Perplexity-User7. And AI Overviews and AI Mode are part of Google Search: they depend on Googlebot, not on Google-Extended11.

A commented robots.txt for WordPress

This example lets every crawler in, AI search crawlers included, except into the dashboard, and opts out of training:

# General rule. Googlebot, Bingbot and the AI search crawlers
# (OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot)
# fall here: they may crawl everything except the dashboard.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

# Model training: opted out (editorial decision).
# Blocking these does not remove you from ChatGPT, Claude,
# Perplexity or Google Search.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
Disallow: /

Sitemap: https://www.example.com/wp-sitemap.xml

Why there is no separate group for the AI search crawlers: Google explains that each crawler obeys only the group with the most specific user agent that matches it12. Create a User-agent: OAI-SearchBot group and that bot stops reading the general group, so you would have to repeat every rule. As long as you want to treat them like Googlebot, the general group is enough.

How to edit robots.txt in WordPress

Pick one route. Combine a physical file, a plugin and a filter, and the result is easily not the one you think.

With Yoast SEO

Yoast's documentation points to Yoast SEO > Tools > File editor. The menu only appears when the file is writable and file editing has not been disabled on your install; otherwise Yoast suggests creating the file over FTP13.

With Rank Math

Switch on Advanced Mode and go to Rank Math SEO > General Settings > Edit robots.txt. Rank Math serves a virtual file, so you must first delete any physical robots.txt3.

With a PHP filter

If you do not use an SEO plugin, or prefer code, a must-use plugin on the robots_txt filter appends the training block to the core file. Save it as wp-content/mu-plugins/robots-ai.php:

<?php
/**
 * Plugin Name: robots.txt for AI crawlers
 * Description: Opts out of model training without touching AI search crawlers.
 */

if ( ! defined( 'ABSPATH' ) ) {
	exit;
}

add_filter( 'robots_txt', function ( $output, $is_public ) {
	// If the site is set to discourage search engines, leave it alone.
	if ( ! $is_public ) {
		return $output;
	}

	$training = array(
		'GPTBot',
		'ClaudeBot',
		'Google-Extended',
		'Applebot-Extended',
		'CCBot',
		'meta-externalagent',
	);

	$output .= "\n# Model training: opted out (editorial decision)\n";
	foreach ( $training as $bot ) {
		$output .= "User-agent: {$bot}\n";
	}
	$output .= "Disallow: /\n";

	return $output;
}, 20, 2 );

Several consecutive User-agent lines form a single group that shares the rules below them12. It has no effect if a physical robots.txt exists.

Do you need an llms.txt on WordPress?

llms.txt is a proposal by Jeremy Howard (September 2024) to publish at /llms.txt a Markdown summary of the site with links to its main pages, meant for language models to read14. It is not a standard and no AI search engine has documented using it. Google says you need no special text file for AI to appear in its AI features11, and in June 2026 it clarified that such files neither help nor hurt in Google Search, although you can keep them for other services15. The evidence on who actually requests the file is gathered in the llms.txt guide.

If you still want one, it is cheap:

  • Yoast SEO generates it from version 25.3 (10 June 2025)16. According to its help pages: Yoast SEO > Settings > Site Features, AI tools section, LLMS.txt option17. You can let Yoast pick the content, giving priority to what you mark as cornerstone, or choose it yourself.
  • Rank Math has an LLMS Txt module that you switch on under Rank Math SEO > Dashboard and configure under General Settings > Edit llms.txt18. Its documentation does not say which version introduced it, and the guide itself describes llms.txt as a proposal, not an official standard18.

Treat it as an optional extra: robots.txt and real crawler access come first.

When robots.txt says yes and the server says no

A robots.txt that allows OAI-SearchBot is worthless if the server answers it with a 403. On WordPress this happens more often than you would think:

  • Security plugins with rules that block "suspicious" user agents or bot lists.
  • The host's firewall, which sometimes blocks crawlers without telling you.
  • The CDN. Cloudflare, for instance, has a setting to block AI bots; the guide to Cloudflare and AI bots walks through it.

A user-agent block on the server is as effective as one in robots.txt, but you will not see it by reading the file. The reverse also holds: blocking by user agent does not protect you from anyone who pretends to be another bot. Anthropic also advises against blocking by IP address, because it can stop its bot from reading your robots.txt6.

How to check it

  1. Open https://www.example.com/robots.txt in a browser and read what is actually served.
  2. Look at the robots.txt report in Search Console: it shows which file Google found, when it last crawled it and any errors, and lets you request a recrawl19.
  3. Request a page with a user agent that contains the bot's name and compare the status code with a browser's. It is an approximation: some firewalls also verify the IP.
curl -sI -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" https://www.example.com/
curl -sI https://www.example.com/

What SmoothSeen does with this

SmoothSeen reads the robots.txt of the page you analyse and separates AI search crawlers from training crawlers: blocking the training ones is not counted as a failure. It also requests the page with the user agent of each AI crawler, a browser and Googlebot, to detect whether the server or firewall answers 403 to any of them, and checks whether an llms.txt exists. All of it feeds the AI visibility score.

What to do this week

Open your /robots.txt, make sure OAI-SearchBot, Claude-SearchBot and PerplexityBot are not blocked, and decide what to do about the training crawlers. Then test with curl that your server does not send them a 403. To have all of that access checked in one go, run an AI visibility audit.

Frequently asked questions

Where is the WordPress robots.txt if I cannot see it over FTP?

Nowhere on disk: WordPress builds it on the fly whenever someone requests /robots.txt and no file with that name sits in the web root. That is why it does not show up in FTP or in your host's file manager. To change it, use an SEO plugin or the robots_txt filter, or create a physical file, knowing that WordPress then stops having a say.

If I block GPTBot, do I disappear from ChatGPT?

No. GPTBot is the crawler OpenAI uses to train models, and its documentation says each setting is independent. What decides whether you appear in ChatGPT search answers is OAI-SearchBot. You can block GPTBot and allow OAI-SearchBot, a common combination for publishers who want to be cited without handing their content over for training.

Does Google-Extended affect AI Overviews or AI Mode?

No. Google-Extended is a robots.txt token that controls whether Google uses your content to train Gemini and for grounding in its apps, and Google states it does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are part of Search and rely on Googlebot; to limit them, Google points to nosnippet and noindex.

Do Yoast SEO or Rank Math create llms.txt automatically?

Only once you switch the feature on. In Yoast SEO it sits in the site features settings, under the AI tools heading, and the free version includes it. In Rank Math it is a module you enable from its dashboard. Either way the file is published at the root of the domain; remember that it is optional and Google does not use it.

Sources

  1. 1WordPress 7.1 source: functions.php (do_robots) and class-wp-sitemaps.php, WordPress (official wordpress-develop repository), accessed 7 October 2026.
  2. 2do_robots(), WordPress Developer Resources, accessed 7 October 2026.
  3. 3Using Rank Math's robots.txt generator, Rank Math, accessed 7 October 2026.
  4. 4Changes to prevent search engines indexing sites, Make WordPress Core, 2 September 2019.
  5. 5Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
  6. 6Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
  7. 7Perplexity Crawlers, Perplexity, accessed 7 October 2026.
  8. 8Google's common crawlers, Google Search Central, updated 14 July 2026.
  9. 9About Applebot, Apple, updated 4 September 2026.
  10. 10CCBot, Common Crawl, accessed 7 October 2026.
  11. 11AI features and your website, Google Search Central, updated 10 December 2025.
  12. 12How Google interprets the robots.txt specification, Google Search Central, updated 31 August 2026.
  13. 13How to edit robots.txt through Yoast SEO, Yoast, accessed 7 October 2026.
  14. 14The /llms.txt file, Jeremy Howard, 3 September 2024.
  15. 15Latest Google Search documentation updates, Google Search Central, updated 1 October 2026.
  16. 16Yoast SEO 25.3 changelog, Yoast, 10 June 2025.
  17. 17How to enable llms.txt with Yoast SEO, Yoast, accessed 7 October 2026.
  18. 18How to use llms.txt in Rank Math SEO, Rank Math, accessed 7 October 2026.
  19. 19robots.txt report, Search Console Help, accessed 7 October 2026.

How to cite this article

SmoothSeen. (2026, October 7). WordPress robots.txt for AI crawlers: allow search, block training, and llms.txt. https://smoothseen.com/en/blog/wordpress-robots-txt-llms-txt/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First version.

Keep reading

WordPress robots.txt for AI crawlers and llms.txt