Skip to content

Perplexity SEO: how Perplexity chooses and cites its sources, and what you can control

Published 7 October 202610 min readBy the SmoothSeen editorial team

Perplexity SEO is the work of getting Perplexity's answer engine to choose your page as one of the numbered sources it links in each answer. It rests on three things you control: that its crawler, PerplexityBot, can read your site, that each section makes sense on its own and that the information is current. Nobody can guarantee the citation.

Key points

  • Perplexity searches the web at the moment of each question and links its sources with numbered citations; its index tracks more than 200 billion URLs, according to the company (September 2025).
  • PerplexityBot is its search crawler, follows robots.txt and is not used to train models. Perplexity-User visits pages at a user's request and, according to Perplexity, generally ignores robots.txt.
  • In August 2025 Cloudflare accused Perplexity of crawling with a disguised browser; Perplexity denied it and blamed a third-party service. Neither account has been independently verified.
  • What you control is PerplexityBot's access (robots.txt and server), sections that stand on their own, up-to-date dates and sourced data. Nobody can guarantee a citation.

To check it on your own site: AI visibility audit

On this page

Perplexity is one of the assistants that AI search optimization has to account for, and the one that shows its sources most openly. This guide brings together what Perplexity itself documents about how it searches and crawls, what was argued in 2025 about how it really does it, and what you can check on your own site this week.

How does Perplexity answer and choose its sources?

Perplexity is an answer engine: it searches the internet at the moment of each question, summarises what it finds with a language model and links each claim to its source with a numbered citation. Its help centre puts it this way: every answer includes numbered citations linking to the original sources, so the user can check the information1. The same article says it works with models from several companies, including OpenAI and Anthropic.

Perplexity has published how its search works under the hood. It did so when it launched its search API, which the company says runs on the same infrastructure as its public answer engine2. The technical article that accompanied the launch gives these figures3:

  • Its index tracks more than 200 billion unique URLs and serves 200 million queries a day.
  • It retrieves candidates with lexical (word-based) and semantic (meaning-based) search at the same time, then ranks them in several stages; the last one uses more powerful reranking models.
  • It splits each document into self-contained spans and scores them one by one. The company's reasoning is that a model answers worse when it is fed irrelevant context.
  • A machine learning model decides which URLs to re-index and when, based on their importance and how often they change.

The practical consequence is direct. Perplexity does not cite your whole page: it cites the passage that answers best. A section that only makes sense once you have read the three before it stands less chance than one that opens with the answer and names its subject. How to write that way is covered in the guide on how to get cited by AI.

Which crawlers does Perplexity use, and what does each one follow?

Perplexity documents two agents, with different jobs and different rules4:

Agent
PerplexityBot
What Perplexity uses it for
Finding websites and linking them in Perplexity's results
Follows robots.txt?
Yes: Perplexity asks to be allowed in robots.txt and says it respects the rules
Trains models?
No, according to Perplexity
Agent
Perplexity-User
What Perplexity uses it for
Visiting a page when a user asks a question that needs it
Follows robots.txt?
Generally not, because a person started the request
Trains models?
No, according to Perplexity

Both publish their user agent and the list of IP addresses they come from, in perplexitybot.json and perplexity-user.json4. Perplexity advises anyone using a web application firewall to combine user-agent matching with IP verification. In its technical article it adds that PerplexityBot honours crawl-rate limits set in robots.txt and slows down when it detects that a site is unavailable3.

There are two takeaways. Blocking PerplexityBot does not protect you from model training, because according to Perplexity it does not train on what it crawls: it only removes you from its results. And blocking Perplexity-User in robots.txt probably achieves nothing, because the company warns that this agent does not usually read it. If you really want to keep it out, the tool is the firewall, not robots.txt. The difference between search and training crawlers across all the assistants is explained in the guide to AI crawlers and robots.txt.

A robots.txt for Perplexity, annotated

One detail that often gets missed: under the robots.txt standard, a crawler that finds a group with its own name obeys that group and stops looking at the User-agent: * group5. If you name PerplexityBot, repeat in its group anything you do not want any crawler to fetch.

# General rule for every crawler
User-agent: *
Disallow: /basket/
Disallow: /my-account/

# PerplexityBot reads ONLY this group, not the one above,
# so the Disallow lines that also apply to it are repeated.
User-agent: PerplexityBot
Allow: /
Disallow: /basket/
Disallow: /my-account/

# Perplexity-User is left out: according to Perplexity it
# generally ignores robots.txt, so a group for it guarantees nothing.

Check what your server does as well

robots.txt is what your site declares; the server may do something else. Some hosting providers and firewalls reject AI crawler user agents by default, and the site owner never finds out. Compare what a browser receives with what a client announcing itself as PerplexityBot receives (replace example.com with your domain):

# 1. What a browser receives: should be 200
curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36" \
  https://www.example.com/

# 2. What PerplexityBot's published user agent receives
curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" \
  https://www.example.com/

If the first command returns 200 and the second 403, something on your server or CDN is rejecting PerplexityBot by name. Be careful with the opposite reading: if your firewall verifies crawler IP addresses, it will reject this imitation and let the real PerplexityBot through, since the real one comes from the addresses on its list. In that case the 403 tells you nothing about the genuine crawler.

What happened between Cloudflare and Perplexity in 2025?

On 4 August 2025, Cloudflare accused Perplexity of using undeclared crawlers to get around websites' blocks, and Perplexity replied that Cloudflare had confused its traffic with another service's. Both sides published figures the other disputes, and we know of no independent verification.

What Cloudflare says6:

  • Customer sites that had blocked PerplexityBot were still seeing their content in Perplexity's answers.
  • Cloudflare registered new, unpublished domains with a robots.txt that disallowed everything. When it asked Perplexity about them, it got detailed information about their content.
  • According to its analysis, when the declared crawler hit a block, requests appeared with a generic Chrome-on-macOS user agent, from changing IP addresses and networks. It put the declared traffic at 20-25 million requests a day and the undeclared traffic at 3-6 million.
  • Cloudflare removed Perplexity from its list of verified bots and added rules to block that crawling. By contrast, it reported that ChatGPT-User fetched the robots.txt and stopped requesting pages when disallowed.

What Perplexity replies7:

  • An assistant that visits a page because a user asked it to is not a crawler, and it compares this with Google's user-triggered fetchers.
  • The 3-6 million daily requests would come from BrowserBase, a third-party cloud browser service that Perplexity says it uses only for specialised tasks, at fewer than 45,000 requests a day.
  • It accuses Cloudflare of not explaining its method and of being unable to tell a legitimate assistant from a threat.

For you, the lesson does not depend on who is right. robots.txt is a request that each company chooses whether to honour, and Perplexity's own documentation already says that Perplexity-User generally does not follow it4. If you need to prevent access, you need a rule on the server or the firewall. If you want to be cited, what matters is that PerplexityBot does not find the door shut.

Does Perplexity have a programme for publishers?

Perplexity has announced two. The first, the Perplexity Publishers' Program, dated 30 July 2024, promised partner media a share of advertising revenue when their content appeared in an answer. It launched with TIME, Der Spiegel, Fortune, Entrepreneur, The Texas Tribune and WordPress.com8.

The second is Comet Plus, announced on 25 August 2025: a subscription giving access to content from partner publishers, whose revenue, minus a portion for compute costs, is shared among them according to three kinds of traffic: human visits, search citations and agent actions9. According to Press Gazette, publishers receive 80% and the first partners were CNN, Condé Nast, Fortune, the Los Angeles Times, The Washington Post, Le Monde and Le Figaro10.

As of this guide, we have not found an official Perplexity page with the current terms or with what has been paid out. What exists are those announcements, and Comet Plus invited interested publishers to get in touch by email. Two caveats: both programmes are aimed at media outlets with commercial agreements, not at any website, and neither announcement says that taking part makes Perplexity cite you more.

What can you control to get cited by Perplexity?

What depends on you fits in this list. None of the points guarantees a citation, but each one removes a reason not to choose you:

What to check
PerplexityBot allowed in robots.txt
Why it matters in Perplexity
It is the agent that links websites in its results
How to check it
Open /robots.txt and look for a group with its name or a blanket Disallow: /
What to check
The server does not reject it
Why it matters in Perplexity
A 403 to its user agent has the same effect as a block
How to check it
The two curl commands above
What to check
Content in the HTML
Why it matters in Perplexity
If the text only appears after JavaScript runs, the crawler may not see it
How to check it
View the page source and search for a sentence from the body
What to check
Sections that stand on their own
Why it matters in Perplexity
Perplexity scores passages, not whole pages
How to check it
Read one section in isolation: does it say what it is about without the rest?
What to check
The answer at the start of each section
Why it matters in Perplexity
The passage that wins is the one that answers first
How to check it
The first sentence under each heading answers that heading
What to check
Visible date and current content
Why it matters in Perplexity
Perplexity re-indexes based on how often each URL changes
How to check it
An updated date on the page and in the markup
What to check
Figures with their source
Why it matters in Perplexity
A figure with a linked origin is easier to cite and to check
How to check it
Every figure carries its source and date in the same sentence

A note on freshness: changing the date without changing anything is not updating. What Perplexity describes is a system that learns how often a page really changes3. If you change the date and the text stays the same, all you lose is credibility with the reader.

What does SmoothSeen check from all this?

Declaration of interest: this blog belongs to SmoothSeen. Its AI ranking analysis checks whether your robots.txt lets PerplexityBot and the other assistants' search crawlers through, and requests the page with each one's user agent to warn you if the server rejects it even though robots.txt allows it. It also checks whether the content depends on JavaScript and whether each section can be cited on its own.

What SmoothSeen does not do is measure whether Perplexity mentions or links to you. It reviews the causes that sit on your page, not the outcome in Perplexity. The questions it tracks on paid plans are in ChatGPT and Gemini; for Perplexity, the way to find out is to ask it your customers' questions yourself, always the same ones, at regular intervals.

Frequently asked questions

Does blocking PerplexityBot stop Perplexity from training on my content?

According to Perplexity, PerplexityBot is not used to train AI models, so blocking it changes nothing in that respect. What it does is stop the crawler from fetching your site to link it in Perplexity's results. If your goal is to stay out of Perplexity, the block makes sense; if you want to be cited, it is exactly the opposite of what you need.

Does Perplexity-User respect robots.txt?

Generally not. Perplexity's documentation says that Perplexity-User acts when a user asks a question and that, because a person started the request, it usually ignores robots.txt rules. If you need to prevent those visits, you will have to block them on the server or the firewall, combining the user agent with the list of IP addresses the company publishes.

How do I know whether Perplexity cites my site?

Ask Perplexity the questions one of your customers would ask and look at the numbered sources in each answer. Repeat the same questions at regular intervals and note who appears, because sources change from one time to the next. In your analytics, visits from its links usually show up as referral traffic from perplexity.ai, when the browser passes on the origin.

Do I need an llms.txt file to appear in Perplexity?

Perplexity does not ask for one in its crawler documentation, nor does it mention it as a signal for choosing sources. What it does document is that its crawler follows robots.txt and that it scores passages from each page. An llms.txt file does no harm, but it does not replace allowing PerplexityBot and writing sections that stand on their own.

What to do next

Open your site's robots.txt, check that PerplexityBot is not blocked and run the two curl commands to see what your server does. Then pick your three most important pages and rewrite the first sentence of each section so that it answers on its own. The other assistants work on the same logic: it is all in the guide to AI search optimization.

Sources

  1. 1How does Perplexity work?, Perplexity Help Center, accessed 7 October 2026.
  2. 2Introducing the Perplexity Search API, Perplexity, 25 September 2025.
  3. 3Architecting and Evaluating an AI-First Search API, Perplexity, 25 September 2025.
  4. 4Perplexity Crawlers, Perplexity, accessed 7 October 2026.
  5. 5RFC 9309: Robots Exclusion Protocol, IETF, September 2022, accessed 7 October 2026.
  6. 6Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, Cloudflare, 4 August 2025.
  7. 7Agents or Bots? Making Sense of AI on the Open Web, Perplexity, 4 August 2025.
  8. 8Introducing the Perplexity Publishers' Program, Perplexity, 30 July 2024.
  9. 9Introducing Comet Plus, Perplexity, 25 August 2025.
  10. 10Perplexity launches AI subscription revenue-share scheme for publishers, Press Gazette, updated 13 November 2025.

How to cite this article

SmoothSeen. (2026, October 7). Perplexity SEO: how Perplexity chooses and cites its sources, and what you can control. https://smoothseen.com/en/blog/perplexity-seo/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First version.

Keep reading

Perplexity SEO: how it chooses and cites its sources