Skip to content

AI visibility audit: a block-by-block checklist to audit a client's site

Published 7 October 20269 min readBy the SmoothSeen editorial team

An AI visibility audit checks whether a page can be read, extracted and cited by assistants such as ChatGPT or Gemini. It covers five blocks: access for search crawlers, answer extractability, credibility (data, sources, authorship and dates), topic coverage, and the requirements specific to the page type.

Key points

  • The audit checks causes in five blocks: access, extractability, credibility, coverage and page-type requirements.
  • Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are not training crawlers (GPTBot, ClaudeBot, Google-Extended), and blocking each has different effects.
  • Google says AI Overviews and AI Mode only need a page to be indexed and eligible for a snippet; there are no extra requirements.
  • Tracking mentions is no substitute for an audit: SISTRIX found 42% to 74% of the domains ChatGPT cites are new each week, depending on the country.
  • Every checklist item comes with a manual test and the evidence to hand to the client.

To check it on your own site: For agencies

On this page

It is the technical and editorial side of AI search optimization: it does not measure whether an assistant names your client today, but whether it has any reason to be able to. What follows is a checklist you can run by hand, block by block, with the evidence worth keeping for the AI section of your SEO client report.

How is auditing causes different from tracking mentions?

Tracking mentions answers "are we cited?"; an audit answers "can we be cited?". They are two tools for two questions, and the second is the one your client can actually fix.

The first answer moves a lot from week to week. According to SISTRIX, between 42% and 74% of the domains cited by ChatGPT Search are new every week, depending on the country (60% in the UK, 74% in Germany), in a study of 82,619 prompts across six countries from 17 December 2025 to 8 April 20261. With that much churn, a single screenshot of an answer proves nothing either way. Causes, by contrast, are stable: a robots.txt that blocks OpenAI's search crawler blocks it every week.

That is why the audit comes first. If the page is not accessible or not extractable, tracking mentions only confirms an absence you could already explain.

Block 1: can the search crawlers get in?

Access is audited per type of bot, not per company. Each provider uses different agents to search, to fetch a page on a user's behalf and to train models, and blocking one is not the same as blocking the others.

Provider
OpenAI
Search and live retrieval
OAI-SearchBot (ChatGPT search results), ChatGPT-User (user actions)
Training
GPTBot
Provider
Anthropic
Search and live retrieval
Claude-SearchBot, Claude-User
Training
ClaudeBot
Provider
Perplexity
Search and live retrieval
PerplexityBot, Perplexity-User
Training
Perplexity says PerplexityBot does not crawl for training models
Provider
Google
Search and live retrieval
Googlebot (AI Overviews and AI Mode draw on the Search index)
Training
Google-Extended (a control token with no user agent of its own)
Provider
Apple
Search and live retrieval
Applebot
Training
Applebot-Extended (it does not crawl either; it is only a permission)

Three nuances to know before you give a verdict:

  • OpenAI is unambiguous: sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, and each agent is set independently, so you can allow OAI-SearchBot and disallow GPTBot2. Changes take around 24 hours to take effect.
  • Google-Extended does not affect inclusion in Google Search and is not a ranking signal. It does control whether content is used to train Gemini and for grounding in Gemini Apps and on Vertex AI3. Blocking it, therefore, does not remove anyone from AI Overviews.
  • User-initiated agents do not always read robots.txt. OpenAI warns that robots.txt rules may not apply to ChatGPT-User2, and Perplexity says Perplexity-User generally ignores them4. Anthropic, on the other hand, says its bots honour robots.txt and that blocking Claude-SearchBot or Claude-User may reduce a site's visibility in its search results5.

Here is a robots.txt that keeps the two decisions apart. It is an example: whether the client wants its content used to train models is the client's call, not a technical fault.

# 1) Search and live retrieval: these are the agents that can link
#    the site in an answer. Blocking them removes you from that search
#    or reduces your visibility in it.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

# 2) Visits a user asks for from the chat. OpenAI and Perplexity warn
#    they may not obey robots.txt; Anthropic says Claude-User does.
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /

# 3) Training: an editorial decision for the client. According to
#    their owners, blocking these does not remove the site from search.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

# 4) CCBot (Common Crawl) builds an open archive of the web.
#    Blocking it does not touch any assistant's search either.
User-agent: CCBot
Disallow: /

# 5) Everything else, including Googlebot, which AI Overviews
#    and AI Mode depend on.
User-agent: *
Allow: /

Access does not end at robots.txt. For AI Overviews and AI Mode, Google asks for the same as for Search: the page must be indexed and eligible to be shown with a snippet, with no additional technical requirements6. A forgotten noindex or nosnippet shuts that door. The text also has to be in the HTML the server sends: in Vercel's December 2024 analysis, none of the major AI crawlers rendered JavaScript, except Gemini (through Googlebot) and Applebot7.

Block 2: can the answer be lifted out without context?

Extractability measures whether a passage makes sense on its own. An assistant quotes fragments, not whole pages, and a fragment that begins "as mentioned above" is no use to it.

What to check:

  • A direct answer up front: an opening paragraph of 40 to 60 words that answers the main query with no preamble.
  • An H1 and H2s that describe what sits below them, ideally phrased as questions when that is what they answer.
  • Lists and tables for steps, comparisons and thresholds: an extractor takes them whole.
  • A visible summary and FAQ, with any markup matching the text.

How to write each of these is covered in the guide on how to get cited by AI. If the client asks about llms.txt, it helps to know that Google explicitly says you do not need new machine-readable files, AI text files or markup to appear in its AI features6; the rest of the evidence is in the article on llms.txt.

Block 3: is there any reason to trust the page?

Credibility is audited by looking for what a third party can verify: figures with sources, a named author, a visible date and an identifiable organisation behind the site.

The best public evidence that this matters comes from the academic paper that proposed the term GEO. In its tests (10,000 queries, presented at KDD 2024), adding quotations, statistics and source citations improved a page's visibility in a GPT-3.5-based generative engine's answers by up to 40%8. It is a lab experiment, not a field guarantee, and that is how you should present it to the client.

Google makes the same point from another angle: it asks whether it is self-evident who wrote the content and whether it provides original information, research or analysis9. In the audit, that becomes five checks: figures with linked sources, a named author with a profile, publication and update dates, an "About us" page with real details, and structured data that does not contradict the visible text.

Block 4: does the page resolve the whole intent?

Coverage measures whether the page answers what the visitor came for, including the follow-up questions. A "boiler replacement cost" page that gives no price ranges, timescales or what is included forces the assistant to find that data elsewhere, and to cite that other site.

To audit it, write down the four or five questions a real buyer would ask and check which ones the page answers with a concrete fact. The ones it does not answer are the gap.

Block 5: does it meet the requirements of its page type?

Each page type has its own requirements, and a generic checklist misses them:

  • Local business: name, address and phone number identical on the site and on the Google Business Profile, opening hours, service area and LocalBusiness markup. Our guide to local SEO for AI search goes into detail.
  • Product page: price, availability, brand and description, both visible and in the Product markup.
  • Category or catalogue page: copy that explains what is on offer and how to choose, not just a grid of products.

Decide the type of each template before you audit. Applying product-page requirements to a blog post produces false findings.

The full checklist

Block
Access
What to check
robots.txt per type of bot
How to check it by hand
Open /robots.txt and look for each agent in the Block 1 table
Evidence for the client
Screenshot of the file with the lines highlighted
Block
Access
What to check
Indexable and snippet-eligible
How to check it by hand
View source for noindex, nosnippet, max-snippet; inspect the URL in Search Console
Evidence for the client
URL Inspection result
Block
Access
What to check
Text in the served HTML
How to check it by hand
curl the URL, or disable JavaScript, and search for the opening paragraph
Evidence for the client
The raw HTML next to the rendered page
Block
Extractability
What to check
Answer in the first paragraph
How to check it by hand
Read the first 60 words without the rest of the page
Evidence for the client
The current paragraph and a rewritten proposal
Block
Extractability
What to check
H1, H2s, lists and tables
How to check it by hand
Review the heading outline
Evidence for the client
The heading outline
Block
Extractability
What to check
Visible summary and FAQ
How to check it by hand
Find them on the page and compare with the markup
Evidence for the client
Schema Markup Validator output
Block
Credibility
What to check
Figures with source and date
How to check it by hand
List every figure and whether it links to its origin
Evidence for the client
List of unsourced figures
Block
Credibility
What to check
Authorship, dates, "About us"
How to check it by hand
Look for a byline, author profile and dates
Evidence for the client
Screenshot of the article header
Block
Coverage
What to check
Intent resolved
How to check it by hand
Answer the buyer's questions using only the page
Evidence for the client
Table of question and answered yes or no
Block
Page type
What to check
Type-specific requirements
How to check it by hand
Compare the site, the Google Business Profile and the markup
Evidence for the client
Discrepancies found

Audit one URL per template (home page, category, product, article, location page), not every URL: template faults repeat on every page that uses the template.

How should you prioritise and present what you find?

Blockers first, everything else after. A block on OAI-SearchBot or a noindex keeps the page out, however good it is; a missing table only makes it less citable. Within each block, order by the number of pages affected: a template fault outweighs a one-off.

For the client, separate three things: what passes, what fails and what could not be checked, with the reason. Nothing you have not checked should appear as fine. And since any deployment can bring back a block or a noindex, the access checks are worth repeating every month; how to organise that is covered in our guide to SEO monitoring.

A declaration of interest: this blog belongs to SmoothSeen, a web audit tool that measures visibility in search engines and AI assistants. Its AI visibility audit groups the checks into four blocks similar to these (extractability; credibility, which includes access for search crawlers; coverage; and page-type requirements), separates search crawlers from training crawlers in robots.txt and flags content that only appears with JavaScript. From the Professional plan, the results come out as a PDF and a link for the client; on Agency, under your brand. For the other question, mentions, paid plans track the questions you set in ChatGPT and Gemini. The checklist above works just as well without SmoothSeen.

Frequently asked questions

What is an AI visibility audit?

It is a review of the causes that determine whether an assistant such as ChatGPT, Gemini or Perplexity could cite a page. It checks whether search crawlers can get in, whether the answer can be extracted without context, whether the page gives reasons to trust it, whether it covers the full intent and whether it meets the requirements of its page type.

How is it different from an SEO audit?

They share a foundation: indexing, speed and structure still matter, and Google uses its normal index for AI Overviews. The AI audit adds access for each assistant's crawlers, the extractability of standalone passages and verifiable credibility, such as sourced figures and visible authorship, which a classic SEO audit often treats as secondary.

How long does an AI visibility audit take?

It depends on the number of templates, not the number of URLs. A site with a home page, categories, product pages and a blog has four templates: one representative URL from each covers most faults, because they repeat on every page sharing that template. Then review separately the pages that bring in the most traffic or revenue.

Should the client block GPTBot to protect its content?

That is the client's decision, not a technical requirement. According to OpenAI, blocking GPTBot signals that content should not be used to train its models, and it is independent of OAI-SearchBot, which decides whether the site appears in ChatGPT search. What you should avoid is blocking both without realising they do different jobs.

What to do next

Open your client's robots.txt and check, agent by agent, whether it lets OAI-SearchBot, Claude-SearchBot and PerplexityBot through. Then run the table against one URL per template and separate the blockers from the improvements. For how the findings fit into your monthly reporting, continue with the SEO client report guide.

Sources

  1. 1AI Citation drift: How stable are sources in AI search results?, Johannes Beus, SISTRIX, updated 12 May 2026.
  2. 2Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
  3. 3Google's common crawlers, Google Search Central, updated 14 July 2026.
  4. 4Perplexity Crawlers, Perplexity, accessed 7 October 2026.
  5. 5Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
  6. 6AI features and your website, Google Search Central, updated 10 December 2025.
  7. 7The rise of the AI crawler, Vercel, 17 December 2024.
  8. 8GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (version 3, 28 June 2024).
  9. 9Creating helpful, reliable, people-first content, Google Search Central, updated 5 October 2026.

How to cite this article

SmoothSeen. (2026, October 7). AI visibility audit: a block-by-block checklist to audit a client's site. https://smoothseen.com/en/blog/ai-visibility-audit/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First published version, sources checked.

Keep reading

AI visibility audit: a block-by-block checklist