Part of the get cited by AI series

Perplexity shows its sources. Here is how to become one of them.

Perplexity lists its sources openly, numbered, under every answer. Getting in takes three things: your page has to be reachable by its crawler, it has to answer the question in its opening lines, and it has to back its claims with something a reader can check. It is the most citation-transparent AI engine there is, which makes it the best place to find out whether your content is genuinely quotable.

Check if Perplexity cites me

3 minutes, no credit card

How Perplexity retrieves and ranks its sources

Perplexity works in two stages: it searches the web first, then writes an answer from the pages it has just read. Selection therefore happens in the first stage, before the model writes a single word. A page that does not appear in search results has no route into the answer, however good its content is.

The engine usually rewrites the user question into several distinct searches, pulls back a shortlist of pages, extracts the passages that answer, then assembles a summary with a numbered link to each source. Those passages decide everything. The engine does not cite a page, it cites a fragment of a page. You are not picked because your article is thorough, you are picked because one of its paragraphs answers exactly what was asked.

The practical consequence is uncomfortable: a long, well-structured page can go permanently uncited if none of its paragraphs stands on its own, while a shorter page with sharper phrasing gets pulled into answers week after week.

PerplexityBot and your robots.txt

Perplexity crawls the web with its own agents, PerplexityBot among them, and those agents follow the directives in your robots.txt. Block them and your site cannot be read, so it cannot be cited. This is the most common and most mundane reason a domain is completely absent from Perplexity answers.

The block is rarely deliberate. It arrives through a security module, a web application firewall, an anti-scraping rule added to stop content theft, or a line inherited from an old configuration nobody reads any more. The outcome is identical either way: the crawler leaves empty-handed.

What to check, concretely

  • Open your robots.txt and look for any Disallow rule targeting PerplexityBot or all user agents.
  • Confirm your firewall or CDN is not returning an error to crawlers it does not recognise.
  • Make sure the main content is present in the served HTML, not injected later by JavaScript.
  • Check that your important pages sit in front of any signup wall, form or blocking banner.

Allowing the crawl is an editorial decision, not a technical default. A publisher whose revenue depends on people reading the article itself can reasonably say no. A company looking for visibility has every reason to open the door.

Why Perplexity favours pages that answer directly and cite their own sources

The engine puts its sources on screen next to its answer. It picks the pages it can defend in front of the reader.

The page has to be findable first

Perplexity runs a web search before it writes anything. A page that surfaces nowhere in ordinary search never enters the candidate set, and no amount of rewriting fixes that until the visibility problem is fixed.

The passage has to stand alone

The engine splits your page and quotes a fragment of it. A paragraph that opens with “as we saw above” loses its meaning the moment it is lifted. A paragraph that states a complete fact stays quotable anywhere.

The claim has to be checkable

Perplexity displays its sources next to its answer, so it is exposed. A dated figure, a named author, a link to the primary source: each one is a reason to pick you over a page that asserts and never backs anything up.

The date carries more weight here

A large share of Perplexity questions ask about the current state of something. All else equal, a page that shows a real update date beats an identical page that says nothing about when it was written.

A page that cites its own sources gives the engine something to lean on. It supplies a verification path, which is precisely what a model looks for before repeating a claim as its own.

The role of freshness in Perplexity answers

Freshness counts for more in Perplexity than in an answer built purely from what a model learned during training, because the engine reads the live web at the moment the question is asked. On subjects that move, pricing, product versions, regulation, tool comparisons, an undated page starts at a disadvantage against one that clearly shows when it was last revised.

That is not an argument for mechanical republishing. Bumping a date without touching the content fools nobody for long and costs you trust with readers and engines alike. What works is genuinely revising the parts that have aged, then showing the date of that revision where it can be seen.

On stable subjects the opposite holds. A clean definition, written once and left in place, can keep earning citations for months.

Perplexity, Google and ChatGPT: what actually differs

Three different ways of handling the same question, and three different ways of deciding whether you exist.

Google ranks you

Classic search lines up links and lets the visitor choose. Your position drives your traffic and you read the result in Search Console. The system is legible, documented, and you always know roughly where you stand.

Perplexity numbers you

Every answer carries its source list, visible and clickable. You can see immediately who was kept and on which phrasing. It is the engine where a manual test yields usable information fastest.

ChatGPT cites less often

Depending on the question, ChatGPT answers from memory or triggers a search. Sources appear less consistently, which makes tracking harder to read. The underlying work is the same on both engines.

The rest of the series covers each engine: get cited by ChatGPT, get cited by Claude and appearing in AI Overviews. The shared principles sit on the parent page, get cited by AI.

Measurement

How to check whether Perplexity already cites you

The manual method is immediate: ask Perplexity the questions your buyers actually ask, then read the source list under the answer. If your domain is in it, you are cited. No other engine makes the check this easy, which is why Perplexity so often serves as the observation ground before widening to the rest.

The limit shows up fast. An answer depends on timing, phrasing and conversational context, so a single test proves nothing. You have to repeat the measurement on a fixed set of questions to tell a trend apart from noise.

  • Ask Perplexity the questions your buyers actually ask, not your keywords
  • Note whether your domain appears in the numbered sources, and in what position
  • Look at which specific page is cited, not just the domain name
  • Spot the questions where a competitor is cited instead of you
  • Repeat the measurement on a schedule: the source list shifts week to week

What you read under a Perplexity answer

1. The cited page, with its domain in plain sight
2. The exact passage reused in the answer
3. The other sources kept for the same question
4. The related questions the engine suggests next

BotSEO repeats that check for you on your key questions, engine by engine, and keeps the history so you can see what moved.

What to change on your pages

Move the answer to the top of every section, make each paragraph self-sufficient, then give the reader a way to verify what you claim. Those three moves cover most of what separates an ignored page from a cited one, and they apply to content you already have, without a rewrite from scratch.

Answer inside the first two sentences

A heading that asks a question should be followed immediately by its answer. Development, caveats and examples come after. Six lines of warm-up before the point is a passage the engine will skip.

Write paragraphs that survive on their own

Reread each paragraph as if it were displayed alone, outside the page. If it opens with a pronoun with no referent, or a back-reference to the previous section, rewrite it so it names its own subject.

Cite your own sources

Link the studies, official texts and data you rely on, and date them. You make verification easy, you reduce what can be disputed, and you give the engine a concrete reason to keep you over a page that asserts without reference.

Date what moves, and actually revise it

Show an update date on time-sensitive content, then keep that promise. A genuine revision every few months is worth far more than an automatic timestamp that reflects no change at all.

Name your entities the same way every time

One product written four different ways forces the engine to guess whether it is looking at the same thing. A stable name across the site, repeated identically in your structured data, removes the ambiguity.

What changes in the day to day work

The most visible change is editorial rather than technical: you stop writing to fill a page and start writing to be quoted. Introductions get shorter, vague claims disappear because you cannot source them, and sections start opening with their conclusion instead of building towards it.

The second change is about rhythm. A one-off manual test tells you very little. It is asking the same questions on a schedule that reveals progress, or the arrival of a competitor on a question you assumed was yours.

That is the loop BotSEO automates with its seven agents: each one covers a part of the work, from technical auditing to writing, while citation tracking reports question by question who is cited, on which engine, and since when.

Frequently asked questions about Perplexity citations

Should I allow PerplexityBot in robots.txt?

Yes, if you want visibility inside Perplexity answers: a blocked crawler cannot read your pages, so it cannot cite them. The answer can differ for a publisher whose business depends on people reading the article itself. It is a trade-off between exposure and content protection, not a universal technical rule.

Do I need to rank on Google to be cited by Perplexity?

It is not a formal requirement, but it is a considerable advantage. Perplexity starts from a web search to build its candidate set, so if you surface nowhere you never enter the selection. Solid organic visibility remains the foundation the citation work is built on.

How long before this shows an effect?

It depends on your starting point. A page that already ranks, rewritten to answer in its opening lines, can change status within weeks because the engine is rereading a source it already knows. Building authority for a domain nobody knows yet takes much longer, and it is better to measure the result than to promise a deadline.

Why does Perplexity sometimes cite a lower-ranking competitor?

Because a citation rewards the passage, not the position. If the competitor page contains a paragraph that answers the exact question asked, and yours buries the same information inside a longer development, theirs is the one that gets quoted. It is usually a phrasing problem rather than a substance problem.

Does structured data help?

It helps remove ambiguity about who you are, what you sell and who signs your content. It does not replace clear writing: flawless markup on a page that answers nothing will not earn a single citation. Treat it as a supporting layer, not the main lever.

Is there a way to track citations other than by hand?

There is no official console for generative engines equivalent to Search Console. The only reliable method is to reask your key questions on a schedule and record the sources shown. BotSEO automates that loop across several engines and keeps the history so the movement becomes readable.

Measure before you rewrite

See whether Perplexity already cites your site

Add your site, pick your key questions and find out who is being cited in your place. You can also compare the BotSEO plans before you decide.

Start free

3 minutes, no credit cardSeven agents, search and AI citations