Skip to content

AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google

Published 7 October 202618 min readBy the SmoothSeen editorial team

AI search optimization is the work of getting assistants such as ChatGPT, Gemini, Perplexity and Google's AI Overviews to find your page, understand it and use it as a source. It has two parts: AEO, making the answer easy to lift as it stands, and GEO, earning enough of the model's trust to be cited.

Key points

  • AI search optimization has two halves: AEO (making the answer easy to extract) and GEO (getting the model to choose you as a source).
  • Google says there are no extra requirements for AI Overviews or AI Mode: the page must be indexed and eligible to be shown with a snippet.
  • Blocking GPTBot or ClaudeBot does not remove you from ChatGPT or Claude; the crawlers that decide search are OAI-SearchBot, Claude-SearchBot and Claude-User.
  • According to SISTRIX, between 42% and 74% of the domains ChatGPT cites are new each week, depending on the country (82,619 prompts, December 2025 to April 2026): there is no fixed position.
  • In the GEO study (KDD 2024), citing sources, adding quotations and adding statistics improved visibility by 30% to 40%; keyword stuffing did not help.

To check it on your own site: AI visibility audit

On this page

This guide brings together what the official documentation from Google, OpenAI, Anthropic and Perplexity says, plus studies with a published methodology. Where the evidence is thin, it says so. A declaration of interest: this blog belongs to SmoothSeen, a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand. When the product comes up, it is to explain what it checks and what it does not.

What is AI search optimization, and what else is it called?

AI search optimization is the set of changes to a website that make it more likely an AI assistant will read it, extract an answer from it and cite it as a source. The same work goes by four or five names, and none of them is a standard.

Name
AEO
Where it comes from
Answer engine optimisation
What it emphasises
That the page gives an answer that can be extracted without context
Name
GEO
Where it comes from
Generative engine optimisation, the term coined in the academic paper by Aggarwal and colleagues (KDD 2024)
What it emphasises
That a generative model chooses you as a source when it writes its answer
Name
AIO, LLMO
Where it comes from
AI optimisation, large language model optimisation
What it emphasises
Commercial labels for the same thing. "AIO" is also used as shorthand for AI Overviews, so it is best avoided
Name
AI SEO, AI search optimization
Where it comes from
Industry and tool vendors
What it emphasises
The most common everyday labels, still not settled

One clarification that saves confusion: in marketing, "geo" also refers to geographic targeting. In this guide, GEO always means generative engine optimisation, and the guide on what GEO is and what the study behind it measured covers it in depth.

The useful distinction is the two halves. AEO is format: a direct answer at the top, clear headings, lists and tables an extractor can take whole. GEO is trust: sourced data, authorship, dates and references that make a model prefer your page over another one saying the same thing. A page can have one without the other, and then it stands less of a chance against a page that has both. The detailed comparison, with what each acronym adds to classic SEO, is in AEO vs GEO vs SEO.

How do ChatGPT, Gemini and Google choose their sources?

Assistants choose sources in two ways: drawing on what the model learnt during training, or searching the web at the moment of the question. Linked citations to specific pages mostly come from the second, and it is the only one your site can influence from one week to the next.

Live search versus the model's memory

A language model answers with what it learnt during training, which has a cut-off date. When the question calls for it, it searches. OpenAI explains that ChatGPT may search the web automatically when a question would benefit from current information and that, when it partners with other search providers, it typically rewrites the user's query into one or more targeted queries that it sends to them1. Responses that use search can include citations to their sources. What is specific to ChatGPT, from its crawlers to how it recommends businesses, is in the guide to ChatGPT SEO.

Gemini works in a similar way. Google describes grounding as "providing content from the Google Search index to the model at prompt time to improve factuality and relevancy" in the Gemini apps2. By that description, for Gemini to ground an answer in your site, your site has to be in Google's index.

The practical consequence is that the work that moves citations in the short term is live search work: letting search crawlers read you, getting your page indexed and making the answer extractable. What a model learns about you in its next training run is a separate matter, slower and harder to control; the guide to LLM SEO separates the two routes.

Query fan-out in AI Overviews and AI Mode

Google calls the technique that AI Overviews and AI Mode may use "query fan-out": issuing multiple related searches across subtopics and data sources from a single question. According to Google, this lets them display a wider and more diverse set of links than a classic web search3.

In practice, your page can make it into an answer through one specific sub-question, even if it is not in the top ten results for the main query. A section that answers "how much does it cost?" or "what is the difference between A and B?" well has its own chance of being cited. That is why headings phrased as real questions, with short answers directly beneath them, pay off. What Google documents about both features is gathered in the guide to Google AI Overviews and AI Mode.

Why the sources change every week

Cited sources rotate a lot in ChatGPT and in AI Mode; Google's AI Overviews are far more stable. SISTRIX tracked 82,619 prompts in six countries for 17 weeks (17 December 2025 to 8 April 2026) and measured what share of the domains cited each week had not appeared the week before4:

Platform
ChatGPT Search
Cited domains that are new each week
Between 42% and 74%, depending on the country (60% in the UK, 74% in Germany, 42% in France)
Platform
Google AI Mode
Cited domains that are new each week
Between 54% and 59%, depending on the country
Platform
Google AI Overviews
Cited domains that are new each week
5% (in 53% of prompts no source changed at all)

This makes uncomfortable reading for anyone selling "rankings in ChatGPT": there is no stable position in ChatGPT. Being cited today does not guarantee being cited next week, and not being cited today does not mean you never will be. What can be measured meaningfully is the trend over several weeks, and the causes that depend on your page.

The five blocks that decide whether you get cited

What decides whether an assistant cites you can be sorted into five blocks: access, extractability, credibility, coverage and what is specific to your type of page. The order matters. If access fails, nothing else gets a chance to count.

This is a way of organising the work, not an algorithm anyone has published. In SmoothSeen, the AI visibility part of the analysis groups its checks into four of these blocks (extractability, credibility, coverage and page-type specifics) and shows separately which AI crawlers can get in.

Access: letting search crawlers read you

The costliest mistake is confusing training crawlers with search crawlers. Each company publishes its own, and the official documentation keeps them clearly apart:

Crawler
OAI-SearchBot
Company
OpenAI
What it is for
Surfacing websites in ChatGPT's search features
What happens if you disallow it in robots.txt
You are not shown in ChatGPT search answers, though you can still appear as a navigational link
Crawler
ChatGPT-User
Company
OpenAI
What it is for
Visits triggered by a user in ChatGPT
What happens if you disallow it in robots.txt
OpenAI warns that robots.txt rules may not apply, because a person initiates the action
Crawler
GPTBot
Company
OpenAI
What it is for
Collecting content that may be used to train its models
What happens if you disallow it in robots.txt
Your content is not used for training; the setting is independent of search
Crawler
Claude-SearchBot
Company
Anthropic
What it is for
Indexing content to improve search results
What happens if you disallow it in robots.txt
It may reduce your visibility in answers that use search
Crawler
Claude-User
Company
Anthropic
What it is for
Reading a site when a user asks Claude a question
What happens if you disallow it in robots.txt
Claude cannot retrieve your content to answer that question
Crawler
ClaudeBot
Company
Anthropic
What it is for
Collecting content that may be used to train its models
What happens if you disallow it in robots.txt
Your future content is excluded from its training data
Crawler
PerplexityBot
Company
Perplexity
What it is for
Surfacing and linking websites in Perplexity's results; not used for model training
What happens if you disallow it in robots.txt
Perplexity does not spell it out, but it is its search crawler
Crawler
Perplexity-User
Company
Perplexity
What it is for
Visits triggered by a user in Perplexity
What happens if you disallow it in robots.txt
Perplexity says it generally ignores robots.txt
Crawler
Google-Extended
Company
Google
What it is for
Gemini training and grounding of answers in the Gemini apps
What happens if you disallow it in robots.txt
It does not affect your inclusion in Google Search and is not a ranking signal

The sources for each row are the documentation from OpenAI5, Anthropic6, Perplexity7 and Google2. OpenAI sums it up in one sentence: each setting is independent, and a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot so its content is not used for training5.

Two nuances are worth noting. First, as well as training, Google-Extended controls whether your content is used to ground answers in the Gemini apps, so disallowing it can affect what Gemini cites from you, even though it leaves Google Search untouched2. Second, Google does not offer a separate token for AI Overviews or AI Mode. They draw on the Google Search index, so what counts there is that Googlebot can crawl and index the page.

A robots.txt that lets search crawlers in and keeps training crawlers out looks like this:

# Search and live retrieval: allowed
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Allow: /

# Model training: disallowed (an editorial decision)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

Disallowing Google-Extended carries the Gemini cost just described; remove it from the list if you want Gemini to be able to ground answers in your site. The guide to AI crawlers and robots.txt goes through the crawlers company by company, beyond these three.

Beyond robots.txt, Google points out that its long-standing controls also limit what appears in its AI features: nosnippet, data-nosnippet, max-snippet and noindex3. If you restricted snippets years ago for some other reason, check it again.

There is one less visible barrier: content that only appears once JavaScript runs. In an analysis Vercel published in December 2024, none of the major AI crawlers it measured, including OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot, executed JavaScript; Gemini did, because it uses Googlebot's infrastructure8. That measurement is almost two years old and crawlers change, but the cautious conclusion stands: serve your main text in the HTML. The guide to JavaScript SEO for AI crawlers covers this problem in detail.

SmoothSeen checks a fixed list of AI crawlers in your robots.txt, separates search crawlers from training crawlers and shows the result bot by bot. Disallowing training crawlers is a legitimate editorial decision, and it is presented as one. What it cannot claim is that "all" AI bots are allowed or blocked, because no complete register of them exists.

Extractability (AEO): a direct answer, structure and format

A page is extractable when a model can take a passage from it that answers the question without needing the rest. That depends on three things:

  1. The answer comes first. The first two paragraphs say what it is, how it is done or what it costs, with no "in this article we will look at" introductions. This article opens that way.
  2. The structure is machine-readable. A single H1, H2s and H3s without skipped levels, lists for steps and tables for comparisons. An extractor can take a whole table or list as it stands.
  3. The format helps the answer be found. A short summary at the top, headings that ask the questions people actually ask and a visible FAQ section at the end.

Readability belongs here too: text full of endless sentences is harder to summarise without distorting it. The guide on how to get cited by AI works through each point with before-and-after examples. The llms.txt file, often sold as part of this block, has its own guide in the list at the end: read it before spending time on one.

Credibility (GEO): data, sources, authorship and dates

Between two pages that give the same answer, a model has to choose. The study that gave GEO its name measured what tips that choice. Aggarwal and his co-authors built a benchmark of 10,000 queries (GEO-bench) and tested different wording changes on the source pages9:

  • The three best were citing sources, adding quotations and adding statistics, with relative improvements of 30% to 40% on their visibility metric.
  • Keyword stuffing, the old SEO habit, did not help: it scored 17.8 against 19.5 for the untouched page.
  • Adding quotations and adding statistics also improved visibility when tested on Perplexity, a live engine, not just in the lab.

The study itself warns that the effect of each technique varies by topic, so "up to 40%" is a maximum measured on a benchmark, not a result to expect for your site.

Beyond data and sources, there are signals a model can read even if it cannot verify your experience: who wrote the page, when it was published and updated, whether an "About us" page says who is behind it, and whether structured data identifies the publishing entity. They are cheap to add, and the markup must match what the reader sees. The guide to E-E-A-T author signals develops these signals.

Coverage: actually answering the question

A page covers a topic when it resolves the whole intent behind the question, not just the definition. If someone is looking for how to choose home insurance, a 200-word text defining "home insurance" cannot compete with one that explains what is covered, what is excluded and how to compare quotes.

Coverage is not length. The goal is to leave no obvious question unanswered, not to hit a word count. SmoothSeen, for example, flags a page with very little text or one that does not resolve the informational intent, because in those cases there is rarely anything worth citing.

What is specific to your type of page

Each type of page has data an assistant needs in order to recommend it, which a generic article does not carry:

  • Local business: LocalBusiness markup, a structured address, opening hours, coordinates and a Google Business Profile consistent with the website. Without them, an assistant struggles to answer "is it open now?" or "where is it?". The guide to local SEO for AI search covers this in depth.
  • Product page: price, availability, brand, shipping details and returns policy in the markup, plus an original description and reviews.
  • Catalogue or shop home page: marked-up products, crawlable categories and information on shipping and returns.
  • Article or guide: author, dates, sources and a direct answer at the top, which the previous blocks already cover.

Does GEO replace SEO?

No. GEO is built on top of SEO, not instead of it. Google is blunt about it: there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimisations necessary. To be shown as a supporting link, a page must be indexed and eligible to be shown in Google Search with a snippet3.

ChatGPT is similar by a different route. OpenAI says any public website can appear in ChatGPT search as long as it does not block OAI-SearchBot10, and that it sometimes partners with external search providers, sending them a rewritten version of the question1. A page that cannot be crawled or indexed gets through neither door.

What does change is the emphasis:

Aspect
The goal
Classic SEO
The page appears in the results and gets the click
AI search optimization
The assistant uses the page as a source and cites it
Aspect
What competes
Classic SEO
The whole page against other pages
AI search optimization
A passage (a paragraph, a table, a list) against other passages
Aspect
What carries most weight
Classic SEO
Relevance, links, page experience
AI search optimization
The same, plus a direct answer, sourced data and authorship signals
Aspect
How it is measured
Classic SEO
Rankings, impressions and clicks
AI search optimization
Mentions and citations repeated over time, plus the traffic that arrives from assistants

The last point matters because clicks fall when an AI summary appears. In a Pew Research study of 900 US adults (March 2025, 68,879 searches), users clicked a traditional result in 8% of visits with an AI summary, against 15% without one. Only 1% of visits included a click on a link inside the summary itself11. These are US figures from a single platform, but the takeaway is clear: being cited does not guarantee the visit.

How do you measure AI search visibility?

There are two separate questions, and each needs its own tool: why a page is or is not cited (auditing the causes on the page) and whether it is actually cited (tracking the results in the assistants). Mixing them up produces reports that say a lot and explain little.

Tracking results means asking assistants the questions a customer would ask, with search switched on, and noting whether they name you, whether they link to your site and whom they cite instead. Given the rotation SISTRIX measured, a single query is not enough: you need to repeat it over weeks and look at the trend4. For traffic, OpenAI adds the parameter utm_source=chatgpt.com to links that leave ChatGPT, so those visits can be filtered in your analytics tool10. Google, for its part, includes traffic from its AI features in the Search Console Performance report, under the "Web" search type3. How to set up those filters step by step is in the guide to tracking AI traffic in GA4.

Auditing the causes means going through the page block by block: whether search crawlers can read it, whether the answer is extractable, whether there are data, sources, an author and a date, whether it covers the question and whether it carries what its page type needs. That tells you what to fix, and it is the only part you can change yourself. If you work for clients, the AI visibility audit explains how to run one and how to prioritise what you find.

SmoothSeen focuses mainly on auditing the causes: it reviews the page and does not measure whether Perplexity, Claude, Copilot or AI Overviews cite you. On paid plans it can also track the questions you define in ChatGPT and Gemini and show whether they name your brand.

Where to start: a block-by-block checklist

This checklist condenses the guide into ten checks you can run today on your most important page, in this order. Each row says what to look at and why it matters.

Block
Access
What to check
Your robots.txt does not disallow OAI-SearchBot, Claude-SearchBot, Claude-User or PerplexityBot
Why
These are the crawlers ChatGPT, Claude and Perplexity use to read and cite pages in their search
Block
Access
What to check
The page is indexed in Google, does not restrict snippets with nosnippet or a very low max-snippet, and its main text is in the HTML
Why
Google requires indexing and a snippet for AI Overviews and AI Mode, and other AI crawlers do not run JavaScript
Block
Extractability
What to check
The first two paragraphs answer the main question directly
Why
That is the passage an assistant can cite without reading the rest
Block
Extractability
What to check
There is one H1, the H2s ask real questions and steps or comparisons sit in lists and tables
Why
An extractor takes whole blocks, and query fan-out looks for answers to sub-questions
Block
Extractability
What to check
There is a short summary at the top and a visible FAQ at the end
Why
They give short, self-contained answers that are easy to extract
Block
Credibility
What to check
Every figure has a linked source, a date and, where relevant, a sample size
Why
Adding statistics was among the three most effective changes in the GEO study
Block
Credibility
What to check
The page links to primary external sources and attributes its quotations
Why
Citing sources and adding quotations were the other two most effective changes
Block
Credibility
What to check
It is clear who wrote it, when it was published and updated, and an "About us" page exists
Why
These are the trust signals a model can read, even if it cannot verify your experience
Block
Coverage
What to check
The page resolves the whole intent, not just the definition
Why
If the obvious question is left unanswered, the assistant has to find the answer on another page
Block
Page type
What to check
Local business: LocalBusiness, address, opening hours. Product page: price, availability, shipping and returns in the markup
Why
This is the data an assistant needs to recommend you

If a row fails, fix it before moving on to the next. A page with the best writing in the world will not appear in ChatGPT search answers if OAI-SearchBot cannot read it.

The guides in this topic

Each guide in the cluster develops one part of this one:

  • AEO vs GEO vs SEO. Adds the comparison between the three acronyms and the work they share.
  • What is GEO (generative engine optimisation). Adds what the KDD 2024 study that coined the term measured and where its limits lie, so a benchmark is not mistaken for a promise.
  • LLM SEO. Adds the difference between what a model learns in training and what it finds when it searches, and whether mentions away from your site matter.
  • AI crawlers and robots.txt. Adds the list of crawlers by company and separates the ones that decide search from the ones that collect training data.
  • JavaScript SEO for AI crawlers. Adds why your main text needs to be in the HTML crawlers receive, not only after JavaScript runs.
  • llms.txt: what it is, how to create one and what the data says about whether it works. Adds the official documentation and the studies on whether any AI engine actually reads the file, before you spend time creating one.
  • How to get cited by AI. Adds before-and-after examples of direct answers, citable data and sources: extractability and credibility put into practice.
  • E-E-A-T author signals. Adds the authorship, date and "About us" signals that help a model trust your page.
  • Google AI Overviews and AI Mode. Adds what Google documents about its two AI features and the controls that limit what they show from your site.
  • ChatGPT SEO. Adds how ChatGPT decides which sites to cite, which OpenAI crawler must be able to get in and how to tell whether it names you.
  • Perplexity SEO. Adds how Perplexity chooses and cites sources, and the roles of PerplexityBot and Perplexity-User.

Frequently asked questions

What is SEO for AI called?

There is no single name. The most common are AEO (answer engine optimisation) and GEO (generative engine optimisation), the latter coined in an academic paper presented at KDD 2024. AIO, LLMO, AI SEO and AI search optimization are also in circulation. They all describe the same work: getting an AI assistant to find your page, extract an answer from it and cite it as a source.

What is the difference between AEO and GEO?

AEO deals with format: a direct answer at the top of the page, clear headings, and lists and tables an extractor can take as they stand. GEO deals with trust: sourced data, external references, authorship and dates that make a model prefer your page over another one saying the same thing. In practice they are worked on together, because an extractable answer without credibility loses to one that has both.

Does AI search optimization replace SEO?

No. Google says there are no additional requirements to appear in AI Overviews or AI Mode: the page must be indexed and eligible to be shown with a snippet. ChatGPT, for its part, sometimes partners with external search providers and needs OAI-SearchBot to be able to read you. Without crawling and indexing there is no citation. AI search optimization adds emphasis on direct answers, sourced data and authorship, but it starts from sound technical SEO.

How long does it take to appear in ChatGPT?

Nobody can guarantee a timescale. ChatGPT decides in each answer whether to search and which sources to cite, and according to SISTRIX between 42% and 74% of the domains it cites are new each week, depending on the country4. What does depend on you is removing the obstacles: allowing OAI-SearchBot, getting the page indexed and offering an extractable answer with data and sources. From there, track the trend over several weeks instead of relying on a one-off query.

Can you pay to appear in ChatGPT's answers?

Not in the answer itself. OpenAI started testing ads in ChatGPT in the United States on 9 February 2026, on the Free and Go plans, and is gradually extending them to other regions. It states that ads do not influence answers: they run on separate systems, appear below the response and are labelled as sponsored12. Being cited still depends on your page.

What to do next

Pick the page that matters most to you and run it through the checklist above, starting with robots.txt: it is the check that rules out a serious problem fastest. If you want those blocks reviewed on your own site, with what fails ranked by severity, analyse the page with SmoothSeen.

Sources

  1. 1Searching the web with ChatGPT, OpenAI Help Center, accessed 7 October 2026.
  2. 2Google's common crawlers: Google-Extended, Google Search Central, updated 14 July 2026.
  3. 3AI features and your website, Google Search Central, updated 10 December 2025.
  4. 4AI Citation drift: How stable are sources in AI search results?, Johannes Beus, SISTRIX, updated 12 May 2026 (82,619 prompts and 1,548,213 snapshots in six countries, 17 December 2025 to 8 April 2026).
  5. 5Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
  6. 6Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic (Claude Help Center), updated 7 April 2026.
  7. 7Perplexity Crawlers, Perplexity, accessed 7 October 2026.
  8. 8The rise of the AI crawler, Vercel, published 17 December 2024, accessed 7 October 2026.
  9. 9GEO: Generative Engine Optimization, Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, KDD 2024 (GEO-bench, 10,000 queries), accessed 7 October 2026.
  10. 10Publishers and Developers - FAQ, OpenAI Help Center, accessed 7 October 2026.
  11. 11Google users are less likely to click on links when an AI summary appears in the results, Pew Research Center, published 22 July 2025 (900 US adults, March 2025).
  12. 12Ads in ChatGPT, OpenAI Help Center, accessed 7 October 2026.

How to cite this article

SmoothSeen. (2026, October 7). AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google. https://smoothseen.com/en/blog/ai-search-optimization/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First published version, sources checked.

Keep reading

AI search optimization: a guide to AEO and GEO