Blog

How do AI assistants decide which sources to cite?

7 min read

Founder of delegAIte, creator of Complete Health Dentistry, 30 years working with more than 6,000 practice owners and founders

Short answer

An assistant rewrites the question into several searches, pulls pages from a search index, then writes an answer and links the pages that support it. To be cited, a page must be crawlable, indexed and retrieved first. Research says pages with sourced statistics, quotations and citations get quoted more, and third-party sources get cited far more than brand sites.

Key takeaways

  • Citation is the last of four gates: you must be crawlable, indexed, retrieved for the rewritten query, and then chosen as support.
  • Google and OpenAI both describe rewriting one question into several searches, which Google calls query fan-out.
  • Adding sourced statistics, quotations and citations raised visibility in generative answers by roughly 30 to 40% in a peer-reviewed study. Keyword stuffing did not help.
  • A 2025 preprint found AI search cites third-party “earned” sources far more than brand-owned pages, a much stronger tilt than Google’s results show.

Practice owners usually ask this question after they have been left out. They have a better website than the practice down the street, and the assistant cited the other one anyway. The answer is rarely about the website alone, and the mechanics explain why.

How does an AI assistant find sources in the first place?

It searches. Google says AI Overviews and AI Mode use retrieval-augmented generation: its ranking systems pull relevant pages from the Search index, and the model writes from them. OpenAI says ChatGPT rewrites a question into one or more targeted queries and sends them to search providers. The answer is written from what those searches return.

Google calls the rewriting step query fan-out: one question becomes a set of related searches running at once. OpenAI gives a similar example, where a broad question turns into a specific search, then a second, narrower one after the first results come back. The practical meaning for a practice is that the assistant may never search the words your patient typed. It searches for the sub-questions it thinks it needs answered.

What has to be true before a page can be cited?

Four gates, in order. The page has to be reachable by the assistant’s crawler, present in the index it searches, retrieved for one of the rewritten queries, and then chosen as the page that supports a sentence. Most practices that never get cited are failing the first or second gate, which no amount of writing fixes.

The four gates between a practice page and a citation (scroll sideways for the rest)
GateWhat decides itWhat a practice controls
1. Reachablerobots.txt, host bot protection, the assistant’s own crawlerAllow OAI-SearchBot, PerplexityBot, Googlebot, Bingbot
2. IndexedGoogle’s index, and the providers ChatGPT searchesIndexing in Google and Bing, a page eligible for a snippet
3. RetrievedRelevance to the rewritten sub-queriesPages that answer specific questions in plain words
4. CitedWhether the page supports a sentence the model wroteClear, sourced, quotable statements that agree with other sites

The vendors document the first two gates plainly. Google says a page must be indexed and eligible to show with a snippet to appear in its generative AI features. OpenAI says a site that blocks OAI-SearchBot will not be shown in ChatGPT search answers, though it can still appear as a plain link. Perplexity says PerplexityBot exists to surface and link websites in its results.

What kind of content gets cited most?

Content that gives the model something specific to repeat. The peer-reviewed GEO study at KDD 2024 tested content changes on a 10,000-query benchmark. Citing sources, adding quotations and adding statistics raised visibility in generative answers by roughly 30 to 40%. Keyword stuffing produced little or no improvement.

One result from that paper should interest smaller practices. The gains were largest for pages that ranked lower. With the cite-sources change, a site ranked fifth in the search results gained 115% in visibility while the top-ranked site lost 30%. Being the biggest name in town doesn’t lock the answer. Being the clearest, best-sourced page on a narrow question can move you into it.

Microsoft gives similar advice to publishers in its Bing Webmaster Tools guidance: clear headings, tables and FAQ sections, claims backed by evidence and data, and pages kept up to date. That is a search company telling you what its assistant finds easy to cite, and it lines up with the research.

Do assistants prefer other websites over the practice’s own site?

Often, yes. A 2025 study by Chen and colleagues compared AI search citations with Google results and found a systematic, strong bias toward earned media, meaning independent third-party sources, over brand-owned pages and social content. Google’s own results showed a more balanced mix. That paper is a preprint, not yet peer reviewed.

For a practice, earned media means the directories, local news, hospital and association listings, and review platforms that talk about you without you writing the words. It explains the complaint I hear most. Your site can be excellent and still be the second source an assistant reaches for, behind a directory that describes you in one line. That line had better be right.

Does ranking on Google mean an assistant will cite you?

Less than you would expect. Ahrefs compared citations from ChatGPT, Gemini, Copilot and Perplexity across 15,000 long-tail queries and found only 12% of links cited by ChatGPT, Gemini and Copilot were in Google’s top ten for the same prompt. Perplexity overlapped more, at about 28.6%. That is a 2025 vendor study.

Some of the gap comes from fan-out. The assistant is searching sub-questions you never ranked for, and citing whichever page answers each one. The rest comes from selection. A page that ranks well on a broad term but never states a specific fact in a sentence can be retrieved and still not be cited.

What can a practice actually influence?

Gates one and two are yours to fix outright, and they are the cheapest. Gate three improves when each important page answers one question a patient would type, in the words a patient uses. Gate four improves when your pages carry specific, sourced statements and when independent sites say the same things about you that you say about yourself. None of it is guaranteed. OpenAI says so directly: placement is not guaranteed. What you can do is stop failing the gates you control.

You can watch part of gate four from the outside. Bing Webmaster Tools now reports how often Microsoft Copilot and Bing’s AI summaries cite your pages, which URLs they cite, and the grounding queries the AI used to find them. It covers one family of assistants, and it is the clearest view of citation any vendor offers today.

Questions people ask

How does ChatGPT choose which websites to cite?

OpenAI says ChatGPT rewrites your question into targeted searches, sends them to search providers, and ranks results using multiple factors meant to surface relevant, reliable information. Your site must allow OAI-SearchBot to be eligible. OpenAI does not publish the ranking factors and says placement is not guaranteed.

What is query fan-out?

It is Google’s name for an assistant turning one question into several related searches that run at the same time. Each sub-search can retrieve different pages, which is why an AI answer often cites pages that do not rank for the original question.

Do statistics and quotes really help a page get cited?

In the peer-reviewed GEO study at KDD 2024, adding cited sources, quotations and statistics raised visibility in generative answers by roughly 30 to 40%, the best of the methods tested. The figures must be real and sourced. Invented numbers are a liability, especially in healthcare.

Why does an assistant cite a directory instead of my practice website?

A 2025 preprint by Chen and colleagues found AI search strongly prefers third-party earned sources over brand-owned pages. Assistants tend to trust what independent sites say about you. Keep those listings accurate and consistent with your own site.

Can I see when an assistant cites my site?

Partly. Bing Webmaster Tools reports citations in Microsoft Copilot and Bing AI summaries, and Search Console reports impressions in Google’s AI features. ChatGPT and Perplexity publish no equivalent report, so a monthly manual check of fixed questions is still needed.

Sources

  1. Google Search Central, "Optimizing your website for generative AI features on Google Search" (2026): Retrieval-augmented generation, query fan-out, and the indexing and snippet eligibility requirement.
  2. OpenAI Help Center, "Searching the web with ChatGPT" (2026): Query rewriting, use of search providers, ranking by multiple factors, and that placement is not guaranteed.
  3. OpenAI, "Overview of OpenAI crawlers" (2026): OAI-SearchBot surfaces sites in ChatGPT search, and sites that block it are not shown in search answers.
  4. Perplexity, "Perplexity crawlers" documentation (2026): PerplexityBot surfaces and links websites in Perplexity results.
  5. Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024 (peer reviewed): The 30 to 40% gain from sources, quotations and statistics, the keyword stuffing result, and the 115% gain for a fifth-ranked site.
  6. Chen, Wang, Chen and Koudas, "Generative Engine Optimization: How to Dominate AI Search" (2025 preprint, not peer reviewed): AI search shows a strong bias toward earned third-party sources over brand-owned and social content.
  7. Microsoft Bing Webmaster Blog, "Introducing AI Performance in Bing Webmaster Tools" (February 2026): Microsoft’s guidance on structure, evidence and freshness for content cited in AI answers.
  8. Ahrefs, AI search and Google overlap study (2025, 15,000 long-tail queries, vendor study): Only 12% of links cited by ChatGPT, Gemini and Copilot were in Google’s top ten, and about 28.6% for Perplexity.

Published . Last reviewed .

Want this in your business?

Book a call and we will scope exactly which part of your operation an agent should take first.

Book a call with the team