No. 09Insights · AI discovery3 min read

AI discovery

How AI assistants choose which sources to cite

Published
Updated
Reading
3 min read
Topic
AI discovery

How do AI assistants choose which sources to cite?

When AI assistants search the web, they run one or more searches based on the question, retrieve pages they are allowed to crawl, and cite the pages they used to build the answer.

Google describes this as query fan-out. Ranking in the underlying search, crawl access, and whether a page clearly answers part of the question all matter. The exact selection logic is not public, and results vary between runs.

Part of the ai discovery guides. Start with what AEO means for a B2B SaaS company.

What is documented

The AI companies publish only part of how this works. What they do publish is useful.

  • Google. Its documentation for AI Overviews and AI Mode says these features may use a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources", which lets it show "a wider and more diverse set of helpful links" than a classic web search. Pages must be indexed and eligible to show a snippet in Search to be shown as a supporting link. Google's guide to generative AI features adds that these features use retrieval-augmented generation, relying on "core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index."
  • OpenAI. OAI-SearchBot crawls pages for ChatGPT's search features. Sites that block it are not shown in ChatGPT search answers. ChatGPT-User makes requests when a user asks it to visit a page.
  • Perplexity. PerplexityBot indexes sites to surface and link them in Perplexity's answers, and Perplexity-User fetches pages on demand for a user's question.

A simple model of the process

Five-step diagram: the assistant reads the question, writes searches (Google calls this query fan-out), retrieves pages it is allowed to crawl, reads the relevant parts, then answers and cites
A simplified model of how an AI assistant finds and cites sources. Diagram by OwnedSignal.
  1. The assistant interprets the question and decides whether to search.
  2. It generates one or more search queries, often more specific than the user's wording.
  3. It retrieves candidate pages from a search index.
  4. It reads the parts of those pages that relate to the question.
  5. It writes an answer and cites the pages it relied on.

This is a simplification. It helps explain why a page that ranks for a narrow sub-question can be cited even if it does not rank for the original question.

What tends to matter

From the documentation and published studies, these factors appear repeatedly:

  • Crawl access. If the crawler is blocked, the page cannot be used.
  • Search visibility. Pages that rank for the sub-queries are more likely to be retrieved. The relationship is not one to one: because of query fan-out, Google says its AI features can link to a wider set of pages than a classic search for the same question.
  • A clear answer to a specific question. A page that answers one question directly is easier to use than a page that mentions it in passing.
  • Source type. Studies by tracking vendors find assistants cite third-party sources heavily: reviews, forums, video, news and professional profiles. Digiday, using Meltwater data, reported YouTube as the most-cited platform across eight AI products in August 2026.

What is not known

  • The weighting of each factor
  • How much the model's training data influences which sources it trusts
  • How often each system's retrieval changes

Be wary of anyone who claims precise knowledge of the algorithm. The published research is mostly from vendors that sell tracking tools, and its findings shift from month to month.

Engines differ

ChatGPT, Perplexity and Google draw on different indexes and show citations differently. Perplexity shows sources prominently on every answer. Google's AI features sit inside search results. ChatGPT cites when it searches. A company can be well cited in one and missing in another, which is why testing each separately matters. See how to test whether AI assistants know your company.

What to do with this

The practical actions are the same ones that help search in general: allow the relevant crawlers, publish pages that answer specific buyer questions, keep your company information consistent, and be present in the third-party sources your buyers and the assistants trust.

Sources

  1. AI features and your website, Google Search Central, 2025-12-10. Re-read 5 October 2026.
  2. Optimizing your website for generative AI features on Google Search, Google Search Central. Read 5 October 2026.
  3. Overview of OpenAI crawlers, OpenAI.
  4. Perplexity crawlers, Perplexity.
  5. In graphic detail: LLMs keep citing YouTube in search results, Digiday, 2026-09-23. Uses Meltwater citation data.
  6. GEO and AEO: what the evidence supports, Papercrane. Summarises Ahrefs llms.txt data and other studies. Secondary source.

Written by Alex Iliescu, founder of OwnedSignal. I build founder presence for technical AI and SaaS founders.

Next step

Let's look at what a buyer finds when they look you up.

The fit call · 30 minutes

A 30-minute fit call. We look at your profile, your site and one or two AI answers together.

Then I tell you which engagement fits, or whether none does yet.