llms.txt: what it is, how to create one and what the data says about whether it works
Published 7 October 20269 min readBy the SmoothSeen editorial team
llms.txt is a Markdown file, proposed by Jeremy Howard in September 2024, that sits at a site's root and tells language models which content matters. It is not a standard: Google Search ignores it1, and an Ahrefs study of 137,210 domains found that 97% of these files received no requests at all in May 20262.
Key points
- llms.txt is a proposal by Jeremy Howard (3 September 2024), not a standard; version 2 arrived on 10 August 2026.
- Google's generative AI guide says Google Search ignores llms.txt files: they neither help nor harm.
- Ahrefs studied 137,210 domains in May 2026 and found 97% of valid llms.txt files received no requests at all.
- AI bots and agents made only 19.5% of llms.txt requests; AI search crawlers made about 200 between them.
- Writing one takes ten minutes. It makes sense for technical documentation, not as a route to ChatGPT citations.
To check it on your own site: AI visibility audit
On this page
What is llms.txt?
llms.txt is a plain-text Markdown file served at /llms.txt that gives AI agents a summary of a website and links to its key pages. Jeremy Howard published the proposal on llmstxt.org on 3 September 2024, and the site itself describes it as a proposal to standardise, not an agreed standard3.
The format is deliberately minimal. Only one element is required: an H1 with the name of the site or project. After that you can add a blockquote summary, a few paragraphs of context and H2 sections containing lists of links, each followed by a short note after a colon3. The proposal also recommends serving a clean Markdown version of each page at the same URL with .md appended.
Version 2, published on 10 August 2026, shifted the emphasis from stuffing more context into a model to helping agents discover content. It adds a way to advertise the file with a rel="describedby" link in your HTML or in a Link header, and it clarifies that an llms.txt covers the pages beneath its path, so a subfolder can have its own4.
If you are new to the subject, our guide to AI search optimization puts llms.txt in context alongside everything else that matters.
How it differs from robots.txt and sitemap.xml
- File
robots.txt- What it does
- Tells crawlers where they may and may not go
- Who acts on it today
- Google, OpenAI, Anthropic and Perplexity explain how to control their crawlers with it
- File
sitemap.xml- What it does
- Lists indexable URLs so search engines can discover them
- Who acts on it today
- Traditional search engines
- File
llms.txt- What it does
- Summarises the site and points an AI agent to key content
- Who acts on it today
- No AI search engine has documented that it uses it
llmstxt.org draws the line itself: robots.txt is about access, whereas llms.txt is read on demand, when an agent needs information3. llms.txt neither blocks nor permits anything.
Does anything actually read llms.txt?
Almost nothing does, and no AI search engine has said in writing that it does. These are the primary sources available as of 7 October 2026.
Google: you don't need special files
In May 2026 Google published a guide to optimising websites for its generative AI features, last updated on 10 July 2026. It states that you don't need machine-readable files, AI text files or Markdown to appear in Google Search, and that Google Search itself doesn't use them1. It adds that you are free to maintain an llms.txt for other services, because it will neither help nor harm your visibility in Google.
The separate page on AI features and your website makes the same point for AI Overviews and AI Mode: there is no special file or markup to add5.
Ahrefs: 137,210 domains and who really requested the file
Ahrefs analysed traffic to 137,210 domains in May 2026 and found that 97% of valid llms.txt files received no requests whatsoever. The study, written by Louise Linehan and published on 15 June 2026, uses Ahrefs Web Analytics data and discards files that turn out to be error pages or HTML rather than genuine Markdown2.
The headline figures:
- 28% of the domains (around 38,000) publish a valid llms.txt. Ahrefs warns that its customers are more technical than the web at large, so treat that as an upper bound.
- Roughly 22,000 requests reached those files across the whole month. SEO audit tools made 21.7% of them; AI bots and agents together made 19.5%.
- AI search crawlers (OAI-SearchBot, PerplexityBot and Claude's search crawler) made about 200 requests between them, or 1.1%.
- Claude Code, Anthropic's coding agent, fetched more llms.txt files than any AI search crawler or AI assistant.
Ahrefs' verdict is blunt: if your aim is to appear in ChatGPT, Perplexity or AI Overviews, the file is largely decoration. The limits are worth stating too: it covers a single month, and the sample leans towards SEO-aware sites.
OpenAI, Anthropic and Perplexity: they publish one, but their bots follow robots.txt
This is where people get confused. OpenAI, Anthropic and Perplexity each publish an llms.txt on their developer documentation sites. We checked on 7 October 2026: developers.openai.com/llms.txt, platform.claude.com/llms.txt and docs.perplexity.ai/llms.txt all return a 200 status.
Publishing one is not the same as reading everyone else's. OpenAI's crawler documentation covers OAI-SearchBot, GPTBot and ChatGPT-User and how to control them through robots.txt; it mentions llms.txt only as the index of its own docs6. Anthropic's page on ClaudeBot, Claude-User and Claude-SearchBot doesn't mention it at all7. Perplexity describes PerplexityBot and Perplexity-User in robots.txt terms and, like OpenAI, links to its llms.txt only as an index of its own documentation8.
Chrome Lighthouse checks for it, but a missing file doesn't count
Lighthouse, Chrome's auditing tool, includes an llms.txt audit in its agentic browsing category. It only fails the audit when the server returns an error for the file. If the file simply doesn't exist (a 404), the audit is marked not applicable because, according to Chrome's documentation, providing one is optional for now9.
So should you create an llms.txt?
An llms.txt costs very little to create, and its benefit for appearing in AI answers is unproven. The sensible decision depends on the kind of site you run:
- Your situation
- Technical documentation, an API or software
- Recommendation
- Create one
- Why
- Coding agents such as Claude Code do request it, according to Ahrefs
- Your situation
- Company site, online shop or local business
- Recommendation
- Optional, low priority
- Why
- No AI search engine documents reading it
- Your situation
- Shopify store
- Recommendation
- Review the one you already have
- Why
- Shopify has served one by default since May 2026
- Your situation
- Little time for technical SEO
- Recommendation
- Leave it until last
- Why
- Other tasks have far stronger evidence behind them
What you shouldn't do is sell it, or buy it, as the key to being cited by ChatGPT. No published data supports that promise.
How to create an llms.txt in ten minutes
Here is a complete example for a made-up business. It uses example.com, a domain reserved for documentation:
# Riverside Joinery
> Bespoke joinery workshop in Leeds, founded in 1998. We make fitted kitchens, built-in wardrobes and solid timber doors for homeowners across West Yorkshire.
We only work within West Yorkshire. We don't sell flat-pack furniture and we don't deliver nationally.
## Services
- [Fitted kitchens](https://www.example.com/kitchens/): materials, typical lead times and how installation works.
- [Built-in wardrobes](https://www.example.com/wardrobes/): door styles, fittings and guarantee.
- [Solid timber doors](https://www.example.com/doors/): available timbers and exterior finishes.
## Company
- [About us](https://www.example.com/about/): the workshop's history and team.
- [Contact and service area](https://www.example.com/contact/): address, opening hours and the towns we cover.
## Optional
- [Blog](https://www.example.com/blog/): timber care guides.What each part does:
- The H1 is the only mandatory element. Use the business name exactly as you want it cited.
- The blockquote (
>) is the summary. Say what you do, where and for whom in two sentences, without slogans. - The standalone paragraph sets boundaries: what you don't do. An agent can't infer that from your pages.
- Each H2 section groups links in the form
[name](url): note. Link only to pages that exist and load. - The
Optionalsection told agents what they could skip in version 1; version 2 keeps it only as a convention4.
Save it as UTF-8 plain text, upload it to the root of your domain and check the response:
curl -sI https://www.example.com/llms.txt
# Look for "HTTP/2 200" and a text Content-Type, not text/htmlIf you get a 200 but the body is your HTML home page, your server is masking a 404: the file isn't there.
On WordPress
Yoast SEO generates the file in its free version. Go to Yoast SEO, then Settings, Site features, AI tools and LLMS.txt. You can let it choose pages automatically or select them manually. It isn't available on multisite installations10. Without a plugin, upload the file via SFTP or your host's file manager.
On Shopify
Since 28 May 2026, every Shopify store serves a default agents.md, and /llms.txt points to the same content11. To serve your own llms.txt, add a templates/llms.txt.liquid template under Online Store, Themes, Edit code. Without that template, Shopify falls back to your agents.md template and then to its own generated default.
What to do instead if you want AI to read your site
If your goal is to appear in answers from ChatGPT, Claude or Perplexity, these tasks have backing in official documentation:
- Check your robots.txt. Let the search crawlers in: OAI-SearchBot, Claude-SearchBot and PerplexityBot. Blocking training crawlers such as GPTBot or ClaudeBot is a separate decision67. Our guide to AI crawlers and robots.txt goes through them bot by bot.
- Write content that can be quoted: the answer first, figures with sources and clear definitions. Our guide on how to get cited by AI covers this in detail.
- Measure before changing anything. An AI visibility audit shows what is actually holding your site back.
How SmoothSeen treats llms.txt
This blog belongs to SmoothSeen, so here is how the tool handles it. SmoothSeen requests /llms.txt at the root of your domain and counts it as present if it returns a 200 status with some content. It does not look for llms.txt files in subfolders. It is one of the access checks in the AI visibility half of the report.
If the file is missing, the report tells you so. Treat it as what it is, a near-zero-cost good practice with no proven effect on citations, and fix whatever blocks crawlers or stops your content being quoted first.
Frequently asked questions
What is an llms.txt file?
It is a Markdown text file published at the root of a website, at /llms.txt, that gives AI agents a summary of the site and a list of links to its most important pages. Jeremy Howard proposed it in September 2024, and version 2 followed in August 2026. It is an open proposal rather than a standard backed by any standards body.
Does Google use llms.txt?
No. Google's guide to optimising for generative AI features says Google Search does not use llms.txt or similar files, and that having one neither helps nor harms your visibility. Google may crawl the file as it would any other text file on your site, but that doesn't mean it treats it specially. You don't need one for AI Overviews or AI Mode either.
Does llms.txt help you rank in ChatGPT?
There is no evidence that it does. OpenAI does not document that OAI-SearchBot, the crawler behind ChatGPT search, reads other sites' llms.txt files. In the Ahrefs study of May 2026, OAI-SearchBot, PerplexityBot and Claude's search crawler made about 200 llms.txt requests between them, out of roughly 22,000 in total.
What is the difference between llms.txt and robots.txt?
robots.txt is about access: it tells each crawler what it may visit, and Google, OpenAI, Anthropic and Perplexity explain how to use it to control their crawlers. llms.txt permits and blocks nothing; it is a content summary for an agent to consult when it needs one. If you want to control how AI crawlers access your site, robots.txt is the tool to use.
How do I create an llms.txt on WordPress?
The quickest route is Yoast SEO: its free version generates the file once you switch the option on in the site features settings, and it lets you pick the pages by hand. If you would rather write it yourself, create a text file from the template in this article and upload it to the root of your domain via SFTP or your host's file manager.
What to do next
If you publish technical documentation, create your llms.txt today using the template above; otherwise, put it at the bottom of your list. Then spend that time on work with real evidence behind it, starting with our guide to AI search optimization.
Sources
- 1Optimizing your website for generative AI features on Google Search, Google Search Central, updated 10 July 2026.
- 2We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read, Louise Linehan, Ahrefs, published 15 June 2026. Sample: 137,210 domains with traffic in May 2026.
- 3The /llms.txt file, Jeremy Howard, llmstxt.org, updated 10 August 2026.
- 4Changes, llmstxt.org, accessed 7 October 2026.
- 5AI features and your website, Google Search Central, updated 10 December 2025.
- 6Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 7Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
- 8Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 9llms.txt, Chrome for Developers, updated 5 May 2026.
- 10How to enable llms.txt with Yoast SEO, Yoast, accessed 7 October 2026.
- 11Customize /llms.txt, /llms-full.txt and /agents.md, Shopify, published 28 May 2026.
How to cite this article
SmoothSeen. (2026, October 7). llms.txt: what it is, how to create one and what the data says about whether it works. https://smoothseen.com/en/blog/llms-txt/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.