Shopify robots.txt for AI crawlers: robots.txt.liquid, search versus training bots, and llms.txt
Published 7 October 20267 min readBy the SmoothSeen editorial team
On Shopify, AI crawlers are controlled through the theme's robots.txt.liquid template: you keep the loop that prints the default rules and add a group at the end that blocks training crawlers such as GPTBot, ClaudeBot or Google-Extended. Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) need nothing: they follow the general group and can cite you.
Key points
- The robots.txt Shopify generates names no AI crawler, so OAI-SearchBot, Claude-SearchBot and PerplexityBot follow the general group and can read products and collections.
- To block training (GPTBot, ClaudeBot, Google-Extended, CCBot and others) you create the robots.txt.liquid template, keep the default rules loop and add a group at the end.
- Giving an AI search crawler its own group makes it ignore the general group, including the cart, checkout and filter rules.
- Blocking crawlers does not stop the product data Shopify Catalog sends to agentic storefronts such as ChatGPT or Microsoft Copilot; that is managed separately.
- Every Shopify store serves a default /agents.md and /llms.txt shows the same content; since 28 May 2026 both can be customised with theme templates.
To check it on your own site: AI visibility audit
On this page
- What is in Shopify's default robots.txt?
- Which AI crawlers should you let in, and which should you block?
- How do you create robots.txt.liquid in Shopify?
- A template that blocks training and keeps Shopify's rules
- The trap of naming AI search crawlers
- What Shopify's robots.txt does not control
- What about llms.txt on Shopify?
- How to check it
- What SmoothSeen does with this
- What to do this week
- Frequently asked questions
Shopify generates every store's robots.txt and keeps it up to date; the only way to change it is through a theme template. This guide explains what the default file contains, which AI crawlers are worth letting in, how to write the template without losing Shopify's rules, and what happens with llms.txt and agents.md. The rest of the platform's technical SEO is in the guide to Shopify technical SEO, and the full list of crawlers is in the guide to AI crawlers and robots.txt.
What is in Shopify's default robots.txt?
According to Shopify, the default file allows public content with Allow: / and keeps the admin, cart, checkout, accounts and orders out of crawling, along with sorted collections (sort_by) and collections filtered with +, because they duplicate the collection1. We checked it on the public demo store for the Dawn theme on 9 October 2026:
- There are two groups,
User-agent: *andUser-agent: adsbot-google, plus theSitemap:line. - It also blocks collection URLs with two filters at once (
/collections/*filter*&*filter*), theme preview parameters and some internal paths. - At the top there are comments pointing AI agents to
/agents.mdand to Shopify's shopping endpoints. - It names no GPTBot, ClaudeBot, OAI-SearchBot or any other AI crawler.
The consequence of that last point: every AI crawler follows the general group. They can read products, collections, pages and the blog, and they stay out of the cart and checkout.
Which AI crawlers should you let in, and which should you block?
The useful question is not «AI, yes or no» but what each crawler reads your store for. The companies split them by purpose2345:
- Purpose
- Search: citing you in answers
- Crawlers
- OAI-SearchBot (ChatGPT), Claude-SearchBot, PerplexityBot
- If you block them
- You stop appearing as a source in those answers
- Purpose
- User-initiated fetches
- Crawlers
- ChatGPT-User, Claude-User, Perplexity-User
- If you block them
- Some may not follow robots.txt
- Purpose
- Model training
- Crawlers
- GPTBot, ClaudeBot, meta-externalagent, CCBot
- If you block them
- Your content stays out of future training
- Purpose
- Control tokens (they do not crawl)
- Crawlers
- Google-Extended, Applebot-Extended
- If you block them
- You opt out of training Gemini or Apple's models
Google-Extended does not affect your inclusion in Google Search; AI Overviews and AI Mode depend on Googlebot56. Applebot-Extended does not remove your pages from Spotlight, Siri or Safari either7. Blocking training is a legitimate editorial choice: OpenAI makes clear that you can allow OAI-SearchBot while blocking GPTBot2.
How do you create robots.txt.liquid in Shopify?
Shopify's steps1:
- In the Shopify admin, go to Online Store.
- On the published theme, open the actions menu (…) and click Edit code.
- Click Add a new template, choose robots and click Create template.
- Edit the template and save.
Shopify flags this as an unsupported customisation: its support team cannot help with the file, and it warns that incorrect use can result in the loss of all traffic1. Two more details from its documentation: changes are instant, and uploading a theme through the Themes section of the admin does not import robots.txt.liquid (Theme Kit or the command line keep it). The template only supports the robots, group, rule, user_agent, sitemap and request objects8.
A template that blocks training and keeps Shopify's rules
Shopify recommends keeping the loop that prints its default groups and adding your rules, rather than replacing the template with plain text; that way its own rules keep updating81. Here is the template, with the loop copied from the official documentation and a custom group at the end:
{% for group in robots.default_groups %}
{{- group.user_agent }}
{%- for rule in group.rules -%}
{{ rule }}
{%- endfor -%}
{%- if group.sitemap != blank -%}
{{ group.sitemap }}
{%- endif -%}
{% endfor %}
# Training crawlers: blocking them does not remove you from AI answers
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /Several consecutive User-agent lines form a single group, as the standard allows9. How we tested it: the template compiles in strict mode with Liquid 5.14 (Shopify's open-source library), using mock objects that imitate Shopify's. We then ran the output through a matcher that applies Google's rules. Products and collections stay allowed for Googlebot, OAI-SearchBot, Claude-SearchBot and PerplexityBot; the cart, sort orders and URLs with two filters are blocked for everyone; and the six training crawlers are blocked across the store. The mock does not reproduce Shopify's servers: compare your live /robots.txt after publishing.
The trap of naming AI search crawlers
It is tempting to add this «to make sure» ChatGPT can read you:
User-agent: OAI-SearchBot
Allow: /Do not. A crawler obeys only the most specific group that names it and ignores the rest10. With that group, OAI-SearchBot stops reading Shopify's User-agent: * group and is free to crawl the cart, checkout, sorted collections and filter combinations. Without a group of its own it could already read everything public. If you ever need custom rules for an AI search crawler, copy the general group's blocks into its group.
What Shopify's robots.txt does not control
Shopify Catalog. If you sell through agentic storefronts such as ChatGPT or Microsoft Copilot, Shopify sends your product data to those channels through Shopify Catalog, independently of robots.txt. Blocking crawlers does not stop that feed; it is managed in your agentic storefronts settings1.
The network layer. Shopify handles bot management on its own network and does not recommend putting a proxy in front of the store, nor can it support one1. So blocking AI bots in Cloudflare in front of Shopify falls outside what Shopify supports. And remember that robots.txt is advisory: not every crawler follows it, and blocking by user agent does not stop anyone who fakes theirs.
Your own audits. If an SEO crawler you run hits rate-limiting errors, Shopify lets you sign it with Web Bot Auth under Online Store > Preferences, in the Crawler access section, with signatures that expire after at most three months. Shopify notes that search engines and large language models can index your store without signatures11.
What about llms.txt on Shopify?
Every Shopify store serves a default /agents.md, and /llms.txt and /llms-full.txt show the same content. On 28 May 2026 Shopify announced that they can be customised12: you add agents.md.liquid to the theme (it covers all three URLs) or llms.txt.liquid (for /llms.txt only). Those templates run in a restricted context: they only see the request and agents objects, not shop or collections13. Shopify recommends sticking with the managed file unless you have advanced needs.
Before spending time on it: Google does not use llms.txt, and no AI search engine has documented reading it6. It is cheap, but its effect is unproven. The guide to what llms.txt is and whether it works explains why.
How to check it
- The live file: open
https://www.example.com/robots.txtand check that Shopify's groups are still there and yours is at the end. - Before and after: Shopify advises saving a copy of your current
/robots.txtand comparing it with what the template renders8; a template that prints the default groups can include rules the generated file no longer has, such asDisallow: /search1. - Search Console: the robots.txt report shows the version Google has read.
- Response per bot:
curl -I -A "OAI-SearchBot/1.0" https://www.example.com/products/a-productshould return 200, just as it does for a browser.
What SmoothSeen does with this
SmoothSeen reads your robots.txt and tells you which AI crawlers it lets in, keeping search crawlers apart from training ones: blocking training crawlers is not counted as a fault. It also requests the page with each crawler's user agent to see whether the network returns a 403 to any of them while serving the page to a browser, and checks whether there is an llms.txt.
What to do this week
Open your /robots.txt and save a copy. If you decide to block training, create robots.txt.liquid with the template above, publish it and compare the output with your copy. To see which AI crawlers actually reach your store, run an AI audit with SmoothSeen.
Frequently asked questions
Does blocking GPTBot on Shopify remove me from ChatGPT?
No. GPTBot collects content to train OpenAI's models. ChatGPT search uses OAI-SearchBot, which follows the general group in Shopify's robots.txt as long as you do not give it a group of its own. And if ChatGPT is active as one of your agentic storefronts, your products reach it through Shopify Catalog, regardless of robots.txt.
Can I block AI bots with Cloudflare in front of Shopify?
Shopify manages bots on its own network and does not recommend a proxy configuration in front of the store, nor can it support one. The supported way to decide which crawlers get in is robots.txt.liquid. If you use a proxy anyway, check that Googlebot and the AI search crawlers still receive a 200.
How do I go back to Shopify's default robots.txt?
Delete the robots.txt.liquid template from the theme. Shopify goes back to serving the file it generates and maintains, with the current rules. Shopify also points to this when a template carries rules the default file no longer includes, such as blocking /search or /policies/: deleting it is how you return to the current file.
Does Shopify's robots.txt affect Google's AI Overviews?
Only through Googlebot. AI Overviews and AI Mode use the Google Search index, so they depend on Googlebot being able to crawl your pages. Blocking Google-Extended does not remove you from Google Search or from those features: it decides whether your content is used to train Gemini and to ground its answers.
Sources
- 1Editing robots.txt.liquid, Shopify Help Center, accessed 9 October 2026.
- 2Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 3Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic (Claude Help Center), updated 7 April 2026.
- 4Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 5Google's common crawlers, Google Crawling Infrastructure, updated 14 July 2026.
- 6AI features and your website, Google Search Central, updated 10 December 2025.
- 7About Applebot, Apple, updated 4 September 2026.
- 8Customize robots.txt and robots.txt.liquid, Shopify.dev, accessed 7 October 2026.
- 9RFC 9309: Robots Exclusion Protocol, IETF, published September 2022.
- 10How Google interprets the robots.txt specification, Google Search Central, updated 31 August 2026.
- 11Crawling your store, Shopify Help Center, accessed 7 October 2026.
- 12Customize /llms.txt, /llms-full.txt and /agents.md, Shopify developer changelog, 28 May 2026.
- 13agents.md.liquid and llms.txt.liquid, Shopify.dev, accessed 7 October 2026.
How to cite this article
SmoothSeen. (2026, October 7). Shopify robots.txt for AI crawlers: robots.txt.liquid, search versus training bots, and llms.txt. https://smoothseen.com/en/blog/shopify-robots-txt-ai-crawlers/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.