LLM SEO (LLMO): what it is and how language models come to mention your brand
Published 7 October 20269 min readBy the SmoothSeen editorial team
LLM SEO, also called LLMO, is the work of getting large language models to mention your brand or cite your site in their answers. A model knows you in two ways: what it learnt in training, up to a cut-off date, or what it retrieves from the web when it answers. Only the second links to sources and can be worked on soon.
Key points
- LLM SEO (LLMO) is the work of getting language models to mention your brand or cite your site. It is a commercial label with no official definition, and it overlaps with GEO and AEO.
- A model knows you in one of two ways: what it learnt in training, with a cut-off date, or what it retrieves from the web when it answers. Only the second links to sources and can be worked on in the short term.
- OpenAI, Anthropic and Perplexity separate their search crawlers from their training crawlers; Google says there are no extra requirements for its AI features.
- The GEO study (KDD 2024) measured gains of up to 40% in a lab, but with the sources already chosen and without touching the model's memory.
- In an Ahrefs study of 75,000 brands, web mentions correlate with AI visibility at between 0.66 and 0.71.
To check it on your own site: AI visibility audit
On this page
This piece explains the concept from the model's side and the mechanics of those two routes; the block-by-block checklist is in the guide to AI search optimization. A declaration of interest: this blog belongs to SmoothSeen, a website audit tool that measures visibility in search engines and AI assistants.
What is LLM SEO?
LLM SEO is the optimisation of a website, and of a brand's wider presence, so that a large language model (LLM) names or cites it when it answers. An LLM is the kind of model behind ChatGPT, Claude or Gemini. LLMO, short for large language model optimisation, is the commercial label for the same work.
None of the names in circulation is a standard, and they overlap heavily. GEO (generative engine optimisation) was born in an academic paper and focuses on engines that write answers with sources; where it came from and what it measured are covered in what is GEO. AEO focuses on format, on making an answer extractable as it stands. The term-by-term comparison, including classic SEO, is in AEO vs GEO vs SEO.
What sets LLMO apart is the point of view: it looks at the model, including when it answers from memory, not only at the search engine feeding it results. Google, for its part, does not see a new discipline. Its guide to generative AI features in Search names AEO and GEO and concludes that, from Google Search's perspective, optimising for generative AI search is "still SEO", because those features are rooted in its core ranking and quality systems1.
The difference that matters is not the name but the route by which the model reaches your brand.
How does a language model come to mention you?
A language model mentions you because it learnt about you in training or because it finds you by searching while it answers. Each route has its own timescale and its own controls.
Memory: what it learnt in training
An LLM learns patterns from vast amounts of text and then answers with what it has learnt. OpenAI explains that its models are trained on three primary sources: information publicly available on the internet, information accessed through third-party partnerships, and information provided by users, human trainers and researchers. It adds that it filters out material it does not want the model to learn from, such as spam, and that the model does not store copies of what it reads: it adjusts its parameters2.
That memory has a use-by date. Anthropic publishes the cut-off for each model: Claude Opus 5.5, for example, was trained on data up to June 2026, and Anthropic itself warns that a model may not know about events after its cut-off date3.
For a brand, this has three consequences:
- It is slow. Whatever you publish today, if it gets in at all, gets in at the next training run, and you do not know when that will be.
- It does not link. When a model answers from memory, it does not cite a URL. It may name you, but it sends you no visits.
- You can only control it going forward. Blocking the training crawlers (GPTBot, ClaudeBot, Google-Extended) excludes your future content, not what has already been learnt.
Retrieval: what it looks up when it answers
Retrieval is the route you can actually work on. The technique has a name: RAG, or retrieval-augmented generation, a method in which the model consults an external collection of documents before writing. The paper that formalised it (Lewis et al., NeurIPS 2020) describes it as combining "parametric" memory, held in the model's weights, with "non-parametric" memory, an index of documents queried at the time of answering4.
Today's assistants do a version of this with the web. Claude, according to Anthropic, invokes a search tool when a question benefits from current information, grounds its answer in live web content and includes citations so the user can check the sources5. ChatGPT rewrites the question into one or more queries that it sends to search providers6. Gemini is grounded in the Google Search index7, and Google calls the technique its AI features use to retrieve pages from that index RAG, or grounding1.
- When your content gets in
- Memory (training)
- At the next training run, with a cut-off date
- Retrieval (search at answer time)
- On every question, if it finds you
- Whether it cites with a link
- Memory (training)
- No
- Retrieval (search at answer time)
- Yes, when the assistant shows its sources
- What controls access
- Memory (training)
- Training crawlers: GPTBot, ClaudeBot, Google-Extended
- Retrieval (search at answer time)
- Search crawlers: OAI-SearchBot, Claude-SearchBot, PerplexityBot; for Google, Googlebot
- What you can do this week
- Memory (training)
- Almost nothing
- Retrieval (search at answer time)
- Open up access, be indexable and offer an extractable answer
What does each engine document?
Every company publishes how its crawlers work, and in all of them the split between training and search shows up. This is what their official pages said when checked on 7 October 2026:
- Company
- OpenAI
- Training
- GPTBot
- Search and live fetching
- OAI-SearchBot (search) and ChatGPT-User (visits a user asks for)
- What it says about appearing
- To be eligible, allow OAI-SearchBot; placement is not guaranteed68
- Company
- Anthropic
- Training
- ClaudeBot
- Search and live fetching
- Claude-SearchBot (indexing for search) and Claude-User (fetching when a user asks)
- What it says about appearing
- Blocking either search agent may reduce your visibility in Claude's answers9
- Company
- Perplexity
- Training
- PerplexityBot is not used to train foundation models
- Search and live fetching
- PerplexityBot (search results) and Perplexity-User (visits a user asks for)
- What it says about appearing
- Perplexity-User generally ignores robots.txt; a change can take up to 24 hours10
- Company
- Training
- Google-Extended (also controls grounding in Gemini)
- Search and live fetching
- The Google Search index, crawled by Googlebot
- What it says about appearing
- No additional requirements for AI Overviews or AI Mode11
A note on Google: Google-Extended does not affect your inclusion in Google Search, but it does decide whether your content grounds answers in the Gemini apps7. Blocking it can reduce what Gemini cites from you.
For ChatGPT specifically, with an annotated robots.txt and what shapes its recommendations, see the guide to ChatGPT SEO.
What does research prove about the answer stage?
The study that gave GEO its name (Aggarwal et al., KDD 2024) is the most cited reference, and it is worth reading from the model's side. It shows that certain wording changes increase the share of an answer that comes from a source. It proves nothing about how to get onto the list of sources, or about the model's memory12.
The set-up. The authors built GEO-bench, a benchmark of 10,000 queries across 25 domains. They simulated a generative engine that took the top five Google results and asked GPT-3.5 Turbo to write an answer from them. They then rewrote one of those sources using nine techniques. Citing sources, adding quotations and adding statistics improved their main visibility metric by 30% to 40%; keyword stuffing did not help12.
What it does not prove, seen as LLMO:
- The retrieval step was fixed. The five sources for each query were always the same. The techniques redistributed visibility among pages already chosen; none of them got a new page onto the list.
- The engine was not a 2026 assistant. It was GPT-3.5 Turbo inside the authors' own set-up. The Perplexity test used 200 queries and uploaded the sources as files, without letting Perplexity search on its own.
- It says nothing about memory. Everything happens at the answer-with-sources stage; the study does not measure whether a model learns your brand beforehand.
That is why writing techniques come after access and indexing. How to apply them, technique by technique and with examples, is in the guide to content that AI will cite.
Do mentions of your brand elsewhere matter?
Everything points that way, although the evidence is correlational. Ahrefs studied 75,000 brands with a DR above 40 across ChatGPT, Google AI Mode and AI Overviews, and published its results on 12 December 202513:
- Branded web mentions correlated with AI visibility at between 0.66 and 0.71, depending on the platform.
- YouTube mentions (in video titles, transcripts or descriptions) were the most strongly correlated factor, at around 0.737.
- The number of pages on a site had almost no relationship, at around 0.194.
- The three platforms largely mentioned the same brands, with an output overlap correlation of 0.779.
Ahrefs used Spearman correlation and warns that correlation is not causation. Even so, the pattern fits both routes: if many sources talk about you, you are more likely to be in the training data and to be found on third-party pages by a live search. Part of LLM SEO is reputation.
How to work on LLM SEO, in order
This order runs from what blocks everything to what adds on top:
- Access. Check that your robots.txt and firewall let OAI-SearchBot, Claude-SearchBot, Claude-User and PerplexityBot through. Without access, there is no retrieval.
- Indexing. Make sure Google and Bing index your key pages and that you are not restricting snippets with
nosnippet. Google requires indexing and snippet eligibility for its AI features11. - Extractable answer. Put the answer first; then figures with sources, attributed quotations and headings that match real questions. This is the part the GEO study does support.
- Third-party mentions. Build presence in comparisons, trade press, forums and videos. It is the slowest route.
- Training, separately. Decide whether you want your content to train models (GPTBot, ClaudeBot, Google-Extended) as an editorial question. It does not affect ChatGPT or Claude search; Google-Extended does affect what Gemini uses for grounding.
- Measurement. Ask the same questions for several weeks and watch the trend; a single query tells you nothing.
A caveat: creating an llms.txt file is not on the list, because no AI search engine documents that it reads one, and Google states explicitly that you do not need new machine-readable files or AI text files to appear in its Search1.
SmoothSeen checks on your page the access in step 1 and the extractability and sourced data in step 3, along with authorship, coverage and the requirements of each page type; it is the AI visibility half of its analysis. For step 6, from the Professional plan upwards, it puts the questions you define to ChatGPT and Gemini every week and tells you whether they cite you. Nobody, SmoothSeen included, can guarantee that a model will mention you.
Frequently asked questions
What does LLMO mean?
LLMO stands for large language model optimisation. It is a commercial label for the work of getting ChatGPT, Claude, Gemini or other models to mention your brand or cite your site. It is not a standard and has no official definition; in practice it is used almost interchangeably with GEO, LLM SEO and AI search optimization.
Is GEO the same as LLMO?
Almost. GEO (generative engine optimisation) was born in an academic paper presented at KDD 2024 and focuses on engines that generate answers with sources, such as ChatGPT with search or Perplexity. LLMO puts the focus on the model, including when it answers from memory. The practical work is the same: access for search crawlers, extractable answers, sourced data and third-party mentions.
How do I get a model to know my brand if it does not search the web?
There is no direct route. What a model knows without searching comes from its training, which has a cut-off date and which no website can ask to join. All you can do is keep your content public and open to training crawlers, and get other sites to talk about you. Even then, you will not know whether or when you have been included.
Does blocking training crawlers cost me visibility?
In ChatGPT and Claude, not in search: OpenAI and Anthropic keep GPTBot and ClaudeBot separate from their search crawlers. You do reduce the chance that future versions of those models will know you from memory. Google has a nuance: Google-Extended also controls whether Gemini uses your content to ground its answers, although it does not affect Google Search.
What to do next
Ask two or three assistants, with search switched on, the question a customer would ask you, and note whom they cite; ask again in a week. Then review your site's access and indexing with the guide to AI search optimization.
Sources
- 1Optimizing your website for generative AI features on Google Search, Google Search Central, updated 10 July 2026.
- 2How ChatGPT and our foundation models are developed, OpenAI Help Center, accessed 7 October 2026.
- 3How up-to-date is Claude's training data?, Anthropic (Claude Help Center), accessed 7 October 2026.
- 4Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al., NeurIPS 2020, accessed 7 October 2026.
- 5Enable and use web search, Anthropic (Claude Help Center), accessed 7 October 2026.
- 6Searching the web with ChatGPT, OpenAI Help Center, accessed 7 October 2026.
- 7Google's common crawlers: Google-Extended, Google Search Central, updated 14 July 2026.
- 8Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 9Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic (Claude Help Center), updated 7 April 2026.
- 10Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 11AI features and your website, Google Search Central, updated 10 December 2025.
- 12GEO: Generative Engine Optimization, Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, KDD 2024, updated 28 June 2024.
- 13Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews (75k Brands Studied), Louise Linehan, Ahrefs, published 12 December 2025.
How to cite this article
SmoothSeen. (2026, October 7). LLM SEO (LLMO): what it is and how language models come to mention your brand. https://smoothseen.com/en/blog/llm-seo/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.