AI visibility audit: a block-by-block checklist to audit a client's site
Published 7 October 20269 min readBy the SmoothSeen editorial team
An AI visibility audit checks whether a page can be read, extracted and cited by assistants such as ChatGPT or Gemini. It covers five blocks: access for search crawlers, answer extractability, credibility (data, sources, authorship and dates), topic coverage, and the requirements specific to the page type.
Key points
- The audit checks causes in five blocks: access, extractability, credibility, coverage and page-type requirements.
- Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are not training crawlers (GPTBot, ClaudeBot, Google-Extended), and blocking each has different effects.
- Google says AI Overviews and AI Mode only need a page to be indexed and eligible for a snippet; there are no extra requirements.
- Tracking mentions is no substitute for an audit: SISTRIX found 42% to 74% of the domains ChatGPT cites are new each week, depending on the country.
- Every checklist item comes with a manual test and the evidence to hand to the client.
To check it on your own site: For agencies
On this page
- How is auditing causes different from tracking mentions?
- Block 1: can the search crawlers get in?
- Block 2: can the answer be lifted out without context?
- Block 3: is there any reason to trust the page?
- Block 4: does the page resolve the whole intent?
- Block 5: does it meet the requirements of its page type?
- The full checklist
- How should you prioritise and present what you find?
- Frequently asked questions
- What to do next
It is the technical and editorial side of AI search optimization: it does not measure whether an assistant names your client today, but whether it has any reason to be able to. What follows is a checklist you can run by hand, block by block, with the evidence worth keeping for the AI section of your SEO client report.
How is auditing causes different from tracking mentions?
Tracking mentions answers "are we cited?"; an audit answers "can we be cited?". They are two tools for two questions, and the second is the one your client can actually fix.
The first answer moves a lot from week to week. According to SISTRIX, between 42% and 74% of the domains cited by ChatGPT Search are new every week, depending on the country (60% in the UK, 74% in Germany), in a study of 82,619 prompts across six countries from 17 December 2025 to 8 April 20261. With that much churn, a single screenshot of an answer proves nothing either way. Causes, by contrast, are stable: a robots.txt that blocks OpenAI's search crawler blocks it every week.
That is why the audit comes first. If the page is not accessible or not extractable, tracking mentions only confirms an absence you could already explain.
Block 1: can the search crawlers get in?
Access is audited per type of bot, not per company. Each provider uses different agents to search, to fetch a page on a user's behalf and to train models, and blocking one is not the same as blocking the others.
- Provider
- OpenAI
- Search and live retrieval
OAI-SearchBot(ChatGPT search results),ChatGPT-User(user actions)- Training
GPTBot
- Provider
- Anthropic
- Search and live retrieval
Claude-SearchBot,Claude-User- Training
ClaudeBot
- Provider
- Perplexity
- Search and live retrieval
PerplexityBot,Perplexity-User- Training
- Perplexity says PerplexityBot does not crawl for training models
- Provider
- Search and live retrieval
Googlebot(AI Overviews and AI Mode draw on the Search index)- Training
Google-Extended(a control token with no user agent of its own)
- Provider
- Apple
- Search and live retrieval
Applebot- Training
Applebot-Extended(it does not crawl either; it is only a permission)
Three nuances to know before you give a verdict:
- OpenAI is unambiguous: sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, and each agent is set independently, so you can allow OAI-SearchBot and disallow GPTBot2. Changes take around 24 hours to take effect.
- Google-Extended does not affect inclusion in Google Search and is not a ranking signal. It does control whether content is used to train Gemini and for grounding in Gemini Apps and on Vertex AI3. Blocking it, therefore, does not remove anyone from AI Overviews.
- User-initiated agents do not always read robots.txt. OpenAI warns that robots.txt rules may not apply to ChatGPT-User2, and Perplexity says Perplexity-User generally ignores them4. Anthropic, on the other hand, says its bots honour robots.txt and that blocking Claude-SearchBot or Claude-User may reduce a site's visibility in its search results5.
Here is a robots.txt that keeps the two decisions apart. It is an example: whether the client wants its content used to train models is the client's call, not a technical fault.
# 1) Search and live retrieval: these are the agents that can link
# the site in an answer. Blocking them removes you from that search
# or reduces your visibility in it.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
# 2) Visits a user asks for from the chat. OpenAI and Perplexity warn
# they may not obey robots.txt; Anthropic says Claude-User does.
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /
# 3) Training: an editorial decision for the client. According to
# their owners, blocking these does not remove the site from search.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
# 4) CCBot (Common Crawl) builds an open archive of the web.
# Blocking it does not touch any assistant's search either.
User-agent: CCBot
Disallow: /
# 5) Everything else, including Googlebot, which AI Overviews
# and AI Mode depend on.
User-agent: *
Allow: /Access does not end at robots.txt. For AI Overviews and AI Mode, Google asks for the same as for Search: the page must be indexed and eligible to be shown with a snippet, with no additional technical requirements6. A forgotten noindex or nosnippet shuts that door. The text also has to be in the HTML the server sends: in Vercel's December 2024 analysis, none of the major AI crawlers rendered JavaScript, except Gemini (through Googlebot) and Applebot7.
Block 2: can the answer be lifted out without context?
Extractability measures whether a passage makes sense on its own. An assistant quotes fragments, not whole pages, and a fragment that begins "as mentioned above" is no use to it.
What to check:
- A direct answer up front: an opening paragraph of 40 to 60 words that answers the main query with no preamble.
- An H1 and H2s that describe what sits below them, ideally phrased as questions when that is what they answer.
- Lists and tables for steps, comparisons and thresholds: an extractor takes them whole.
- A visible summary and FAQ, with any markup matching the text.
How to write each of these is covered in the guide on how to get cited by AI. If the client asks about llms.txt, it helps to know that Google explicitly says you do not need new machine-readable files, AI text files or markup to appear in its AI features6; the rest of the evidence is in the article on llms.txt.
Block 3: is there any reason to trust the page?
Credibility is audited by looking for what a third party can verify: figures with sources, a named author, a visible date and an identifiable organisation behind the site.
The best public evidence that this matters comes from the academic paper that proposed the term GEO. In its tests (10,000 queries, presented at KDD 2024), adding quotations, statistics and source citations improved a page's visibility in a GPT-3.5-based generative engine's answers by up to 40%8. It is a lab experiment, not a field guarantee, and that is how you should present it to the client.
Google makes the same point from another angle: it asks whether it is self-evident who wrote the content and whether it provides original information, research or analysis9. In the audit, that becomes five checks: figures with linked sources, a named author with a profile, publication and update dates, an "About us" page with real details, and structured data that does not contradict the visible text.
Block 4: does the page resolve the whole intent?
Coverage measures whether the page answers what the visitor came for, including the follow-up questions. A "boiler replacement cost" page that gives no price ranges, timescales or what is included forces the assistant to find that data elsewhere, and to cite that other site.
To audit it, write down the four or five questions a real buyer would ask and check which ones the page answers with a concrete fact. The ones it does not answer are the gap.
Block 5: does it meet the requirements of its page type?
Each page type has its own requirements, and a generic checklist misses them:
- Local business: name, address and phone number identical on the site and on the Google Business Profile, opening hours, service area and
LocalBusinessmarkup. Our guide to local SEO for AI search goes into detail. - Product page: price, availability, brand and description, both visible and in the
Productmarkup. - Category or catalogue page: copy that explains what is on offer and how to choose, not just a grid of products.
Decide the type of each template before you audit. Applying product-page requirements to a blog post produces false findings.
The full checklist
- Block
- Access
- What to check
- robots.txt per type of bot
- How to check it by hand
- Open
/robots.txtand look for each agent in the Block 1 table - Evidence for the client
- Screenshot of the file with the lines highlighted
- Block
- Access
- What to check
- Indexable and snippet-eligible
- How to check it by hand
- View source for
noindex,nosnippet,max-snippet; inspect the URL in Search Console - Evidence for the client
- URL Inspection result
- Block
- Access
- What to check
- Text in the served HTML
- How to check it by hand
curlthe URL, or disable JavaScript, and search for the opening paragraph- Evidence for the client
- The raw HTML next to the rendered page
- Block
- Extractability
- What to check
- Answer in the first paragraph
- How to check it by hand
- Read the first 60 words without the rest of the page
- Evidence for the client
- The current paragraph and a rewritten proposal
- Block
- Extractability
- What to check
- H1, H2s, lists and tables
- How to check it by hand
- Review the heading outline
- Evidence for the client
- The heading outline
- Block
- Extractability
- What to check
- Visible summary and FAQ
- How to check it by hand
- Find them on the page and compare with the markup
- Evidence for the client
- Schema Markup Validator output
- Block
- Credibility
- What to check
- Figures with source and date
- How to check it by hand
- List every figure and whether it links to its origin
- Evidence for the client
- List of unsourced figures
- Block
- Credibility
- What to check
- Authorship, dates, "About us"
- How to check it by hand
- Look for a byline, author profile and dates
- Evidence for the client
- Screenshot of the article header
- Block
- Coverage
- What to check
- Intent resolved
- How to check it by hand
- Answer the buyer's questions using only the page
- Evidence for the client
- Table of question and answered yes or no
- Block
- Page type
- What to check
- Type-specific requirements
- How to check it by hand
- Compare the site, the Google Business Profile and the markup
- Evidence for the client
- Discrepancies found
Audit one URL per template (home page, category, product, article, location page), not every URL: template faults repeat on every page that uses the template.
How should you prioritise and present what you find?
Blockers first, everything else after. A block on OAI-SearchBot or a noindex keeps the page out, however good it is; a missing table only makes it less citable. Within each block, order by the number of pages affected: a template fault outweighs a one-off.
For the client, separate three things: what passes, what fails and what could not be checked, with the reason. Nothing you have not checked should appear as fine. And since any deployment can bring back a block or a noindex, the access checks are worth repeating every month; how to organise that is covered in our guide to SEO monitoring.
A declaration of interest: this blog belongs to SmoothSeen, a web audit tool that measures visibility in search engines and AI assistants. Its AI visibility audit groups the checks into four blocks similar to these (extractability; credibility, which includes access for search crawlers; coverage; and page-type requirements), separates search crawlers from training crawlers in robots.txt and flags content that only appears with JavaScript. From the Professional plan, the results come out as a PDF and a link for the client; on Agency, under your brand. For the other question, mentions, paid plans track the questions you set in ChatGPT and Gemini. The checklist above works just as well without SmoothSeen.
Frequently asked questions
What is an AI visibility audit?
It is a review of the causes that determine whether an assistant such as ChatGPT, Gemini or Perplexity could cite a page. It checks whether search crawlers can get in, whether the answer can be extracted without context, whether the page gives reasons to trust it, whether it covers the full intent and whether it meets the requirements of its page type.
How is it different from an SEO audit?
They share a foundation: indexing, speed and structure still matter, and Google uses its normal index for AI Overviews. The AI audit adds access for each assistant's crawlers, the extractability of standalone passages and verifiable credibility, such as sourced figures and visible authorship, which a classic SEO audit often treats as secondary.
How long does an AI visibility audit take?
It depends on the number of templates, not the number of URLs. A site with a home page, categories, product pages and a blog has four templates: one representative URL from each covers most faults, because they repeat on every page sharing that template. Then review separately the pages that bring in the most traffic or revenue.
Should the client block GPTBot to protect its content?
That is the client's decision, not a technical requirement. According to OpenAI, blocking GPTBot signals that content should not be used to train its models, and it is independent of OAI-SearchBot, which decides whether the site appears in ChatGPT search. What you should avoid is blocking both without realising they do different jobs.
What to do next
Open your client's robots.txt and check, agent by agent, whether it lets OAI-SearchBot, Claude-SearchBot and PerplexityBot through. Then run the table against one URL per template and separate the blockers from the improvements. For how the findings fit into your monthly reporting, continue with the SEO client report guide.
Sources
- 1AI Citation drift: How stable are sources in AI search results?, Johannes Beus, SISTRIX, updated 12 May 2026.
- 2Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 3Google's common crawlers, Google Search Central, updated 14 July 2026.
- 4Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 5Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
- 6AI features and your website, Google Search Central, updated 10 December 2025.
- 7The rise of the AI crawler, Vercel, 17 December 2024.
- 8GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (version 3, 28 June 2024).
- 9Creating helpful, reliable, people-first content, Google Search Central, updated 5 October 2026.
How to cite this article
SmoothSeen. (2026, October 7). AI visibility audit: a block-by-block checklist to audit a client's site. https://smoothseen.com/en/blog/ai-visibility-audit/
Keep reading
SEO client reports in 2026: what to include (AI visibility too) and how to prove results
What an SEO client report should cover in 2026, AI visibility included, how to prove results without inflating numbers, and a section-by-section template.
SEO monitoring: what to watch on a client's site, how often, and how to avoid alert noise
What to monitor on a client's website and how often: indexing, robots.txt, certificates, Core Web Vitals, links and security, with useful alerts and no noise.
White label SEO reports: what they should include and what to tell clients about your tools
What a white label SEO report is, what it should carry under your brand, what to tell clients about your tools and the mistakes that give the vendor away.