Perplexity: the crawler that links, and the one that fetches.
Perplexity documents two user agents. One is designed to surface and link websites in its search results; the other visits a page when a user asks. Its documentation recommends allowing the first.
Being cited in Perplexity means appearing as a linked source in an answer. Perplexity documents the crawler it uses to surface and link websites, PerplexityBot, and recommends allowing it. The page cited here does not say how sources are chosen.
Documentation last checked: October 4, 2026. It is re-checked every quarter, and each check is entered in the change log. These pages can change; the statements below are quoted or paraphrased from the sources linked, as they read on that date.
What is documented
| Statement | Source |
|---|---|
| PerplexityBot is "designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." | Perplexity, crawlers |
Perplexity recommends allowing PerplexityBot in robots.txt and permitting requests from its published IP ranges |
Perplexity, crawlers |
| Perplexity-User "supports user actions within Perplexity": when a user asks a question, it might visit a web page | Perplexity, crawlers |
| Perplexity-User "generally ignores robots.txt rules" | Perplexity, crawlers |
| Perplexity-User is not used for web crawling or to collect content for training | Perplexity, crawlers |
What the cited sources do not cover
- How sources are selected and ordered in an answer.
- Any way to request or guarantee a citation.
- A report of citations for site owners.
Where the documentation is silent, this page says so and does not fill the gap.
Access and controls
The documentation recommends two things: allowing PerplexityBot in robots.txt, and permitting requests from its published IP ranges. A site can do the first and still refuse the requests at the network level.
A test protocol for Perplexity
This is a procedure to follow, not a result. No test of Perplexity is reported on this page.
- Check robots.txt for
PerplexityBot. - Check the firewall or CDN. Are requests from Perplexity's published IP ranges permitted?
- Write entity, category and comparison prompts.
- Run each several times and keep the answer and any sources shown.
- List the cited domains, and note which are yours, which are independent, and which belong to competitors.
- Repeat at a fixed interval.
What to record for each run is listed in what is recorded; how many prompts a reading needs is in how many prompts make a reading. The rules a measurement must meet before it is published on this site are in the research methodology.
What to do with the result
- Never cited, and the crawler is blocked Fix access first, at both levels.
- Cited for your own name, not for your category See competitors recommended by AI.
- The answer cites only your own pages See weak independent sources.
- A cited page states something wrong Correct it at that source: see inaccurate or outdated AI answers.
No one can guarantee a mention, a citation or a position in Perplexity. What can be worked on is access, the clarity of the entity, the evidence on the page and the independent sources. The method is The Authority Architecture System.
Frequently asked questions
Does Perplexity use my content to train models?
Perplexity's documentation states that PerplexityBot is not used to crawl content for AI foundation models, and that Perplexity-User is not used to collect content for training.
I blocked Perplexity in robots.txt and it still fetched a page. Why?
One documented explanation: Perplexity-User, which acts on a user's request, generally ignores robots.txt rules. PerplexityBot is the crawler robots.txt addresses.
Can I ask Perplexity to cite me?
The documentation cited here describes access, not a way to request a citation.
Other platforms
- ChatGPT OAI-SearchBot, the referral parameter, and what OpenAI does not publish.
- Google AI Overviews Eligibility through Google Search, and the snippet controls.
- Google AI Mode Follow-up questions and query fan-out, with the same eligibility as Search.
- Gemini Google-Extended, a control separate from those of Google Search.
- Claude Three Anthropic crawlers with three different effects.
How the engines differ, and why each is measured separately, is set out in AI visibility. Related analyses are collected in Insights.
Understand how search and AI currently interpret your authority.
Request a confidential assessment of your visibility, entity signals, narrative consistency and source authority.
Request a Private AssessmentSelective engagements for leaders and organizations with material reputational stakes.