Sign inCreate a free account
Theme

How AI assistants decide which pages to quote

ChatGPT, Perplexity and Google's AI answers cite sources. The pages they pick have a lot in common, and most of it is within your control.

A growing share of people now get their first answer from an AI assistant rather than a list of ten blue links. Ask ChatGPT, Perplexity, Gemini or Google's AI Overviews a question, and you'll often get a short answer with a few sources underneath it.

For a marketer, those citations are the new front page. So I've been spending a lot of time looking at which pages get quoted and which don't. There's no published rulebook, but the pattern is clear enough to act on.

Assistants quote passages, not pages

This is the single most useful thing to understand. When an assistant cites you, it's usually because one specific passage on your page answered the question well enough to lift.

That changes how you should think about a page. It's less "does this page rank for the keyword" and more "is there a paragraph here that answers the question on its own?"

What quotable pages have in common

The answer comes first

If your page is titled "How long does a domain transfer take?", the first paragraph should say how long. Five to seven days, usually, and here's why. Context, caveats and the history of DNS can come after.

Plenty of pages bury the answer under three paragraphs of introduction because that's what we were taught to write. Assistants, like hurried readers, tend to stop at the first passage that answers.

One idea per paragraph

A paragraph that covers pricing, setup time and support at once is hard to quote for any of them. Short paragraphs that each make one point are easy to lift whole.

Facts with a source and a date

"Most people prefer..." is hard to trust. "In a 2025 survey of 1,200 buyers, 41% said..." is easy to cite. Specific numbers with their origin get picked up far more often than vague claims.

Clear headings that match real questions

Headings phrased the way people actually ask ("How much does it cost?", "Can I cancel anytime?") help both readers and machines find the right passage.

Readable without JavaScript

Many AI crawlers read the raw HTML your server sends and don't run scripts. If your content loads client-side, they may see an empty page. Check with View Source, not Inspect.

Make sure the door is open

None of this helps if the assistant's crawler can't reach your site. A lot of robots.txt files block AI bots without anyone meaning to, through an old blanket rule or a plugin setting. Have a look at yours for lines mentioning GPTBot, ClaudeBot, PerplexityBot or Google-Extended.

Whether to allow them is a genuine business decision, and my colleague Ashwini will go into it in a separate post. Blocking them by accident isn't a decision at all.

A quick self-audit

Pick your three most important pages and ask:

  1. Could someone copy the first paragraph into an answer and have it make sense on its own?
  2. Does every key claim have a number or a source?
  3. Is the content in the raw HTML?
  4. Can the AI crawlers you care about actually reach the page?

Every WebRankPage report now includes an AI & answer readiness section that checks these things for you, including which AI crawlers your robots.txt lets in. It's the part of the report I'd read first if you care about where search is heading.

Free forever · no card

Run it on a page you care about.
See what it says.

Nothing is withheld on the free tier. An account adds the part a single report cannot give you: a record of whether anything you changed actually worked.

  • Every report you run, kept
  • Compare a page over time
  • Three reports a day, not one
  • Every check, same as the paid tiers

Create free account

Free forever · no card required

Already registered? Sign in

  • TLS encrypted
  • Instant setup
  • No card needed