> ## Content Index
> Fetch the complete content index at: https://www.bushletter.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI search cites 4 sources per answer, 37% absent from Google
- URL: https://www.bushletter.com/ai-search-cites-4-sources-per-answer-37-absent-from-google/
- Published: 2026-08-05T23:00:00.000Z
- Updated: 2026-08-05T22:59:59.000Z
- Description: Traditional search engines crawl the web, index pages, and hand users a ranked list of links. LLM-based search engines work differently: an AI agent takes a query, breaks it into sub-queries, retrieves a small set of documents, and synthesises a natural-language answer with in-text citations.
- Author: Editor
- Tags: Marketing, AI, Business, #By Takeshi Mori

![Takeshi Mori](https://res.cloudinary.com/dz77sb7j1/image/upload/v1774262608/bushletter/authors/takeshi-mori.png)

By **Takeshi Mori** · 2026-08-04

TLDR

A study of nearly 56,000 queries found AI assistants cite an average of 4.3 URLs per response, against 10.3 for traditional search engines. More than a third of the domains AI tools cite do not appear in Google or Bing results at all, meaning the rules for getting named in AI search are fundamentally different.

KEY TAKEAWAYS

01AI search engines averaged 4.3 URLs per response versus 10.3 for Google and Bing across 55,936 queries.

0237% of domains cited by AI assistants were entirely absent from traditional search engine results.

03Grok omitted external citations in 82% of responses; Gemini did so in 38% of responses.

04Fewer than ten distinct URLs appeared in 80% of all AI search responses studied.

05Despite citing fewer sources overall, AI engines spread citations more evenly across domains than traditional search.

## Why the source pool shrinks to a handful

Traditional search engines crawl the web, index pages, and hand users a ranked list of links. LLM-based search engines work differently: an AI agent takes a query, breaks it into sub-queries, retrieves a small set of documents, and synthesises a natural-language answer with in-text citations.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) The mechanical result is a much shorter citation list, because the model is selecting the few sources it needs to write one answer, not ranking thousands of pages for a human to browse.

That compression is the starting point for a new academic study covering six LLM-based search engines: ChatGPT, Gemini, Perplexity, Grok, Google AI Mode and Copilot, alongside Google and Bing as traditional benchmarks.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) The researchers ran 55,936 queries across all eight systems, logged every URL returned, and compared the two pools.

## Four links, not ten

On average, LLM-based search engines embedded 4.3 distinct URLs and 3.4 distinct domains per response, compared with 10.3 URLs and 7.3 domains for traditional search engines.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) That is less than half the citation volume, and for any business trying to appear in search results, the competitive surface has shrunk sharply.

The compression is not uniform. Grok omitted external citations entirely in 82% of its responses, and Gemini did so in 38% of its responses.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) Perplexity sat closer to the citation-heavy end of the LLM range. That variation matters for any publisher or brand deciding which AI platform to prioritise.

Fewer than ten distinct URLs appeared in 80% of all LLM search responses studied.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) In practical terms, this is a very small room with a very short guest list.

## The 37% gap

The study's sharpest finding is not the volume drop. It is where AI engines are sourcing those few citations. Lead author Peixian Zhang said the results showed that "LLM-SEs return far fewer URLs per response (4.3 vs. 10 on average), yet 37% of the domains are absent from TSEs."[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) Co-author Qiming Ye said: "LLM-SEs cite domain resources with greater diversity than TSEs. Indeed, 37% of domains are unique to LLM-SEs."[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com)

Overall, 37% of the domains cited by LLM-based search engines did not appear in traditional search engine results at all.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) More than one in three of the domains getting named in AI answers would be invisible to anyone monitoring only their Google rankings. The two visibility landscapes are operating on different rules.

There is a structural explanation for the divergence. Traditional search engines optimise for authority signals: backlink profiles, domain age, E-E-A-T scoring, click-through rates. LLM-based engines retrieve documents that are linguistically useful for constructing an answer, which is not the same criterion. A niche trade publication, a forum thread, a specialist database or a government report can satisfy the retrieval step without ever accumulating the link equity that drives Google rankings.

## Same question, different answer

The small source pool creates a second instability. Because so few URLs are cited per response, minor variation in retrieval, such as a slightly different sub-query expansion or a different freshness threshold, can produce a completely different set of named sources for the same question asked twice.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) A company that appears in one run of a query may not appear in the next.

In traditional search, a page might shift a few positions up or down. In AI search, the outcome is binary: you are in the answer or you are not. With citation lists this short, there is no position twelve that still draws a handful of clicks.

The study's Gini index data reinforces this. LLM-based search engines show lower Gini indices than traditional engines, meaning citations are spread more evenly across domains rather than concentrated in a few high-authority sites.[\[1\]](https://arxiv.org/html/2512.09483?ref=bushletter.com) The distribution is flatter, which sounds egalitarian, but the practical effect is unpredictability: no single domain dominates the way Google's top results do.

## Third-party coverage, not your own website

In traditional SEO, the primary lever is your own website: its technical health, its content quality, its backlink profile. In AI search, your website is one of only 4.3 sources that might get cited. The 37% divergence figure tells you that a substantial portion of those citations come from independent third-party sites retrieved because the content there was linguistically useful for building an answer.

What independent sites say about a company, including reviews, comparisons, editorial mentions and forum discussions, is not a secondary signal in this environment. It may be the primary one. If the documents an AI engine retrieves when someone asks about your product category come from industry publications and review aggregators rather than your homepage, then whether you get named is largely decided upstream, by coverage you do not control.

The study analysed six AI engines at a single point in time, and all of these systems are updating their retrieval logic continuously. Grok's 82% no-citation rate may shift as the product matures. Google AI Mode, sitting on top of the world's largest index, may behave differently over time than Perplexity, which pulls from a smaller crawl. The structural insight holds even as the specific numbers move: AI search selects from a very small pool, that pool diverges sharply from Google's results, and the selection criteria reward being talked about independently. For publishers and brands built around Google visibility, the study is a prompt to map a second landscape before its habits harden.

SOURCES & CITATIONS

1. [A Study of LLM-based Search Engines: Citation Behaviour and Domain Coverage (arXiv:2512.09483)](https://arxiv.org/html/2512.09483?ref=bushletter.com)
2. [Ahrefs, AI search and Google ranking overlap study](https://ahrefs.com/blog/ai-search-overlap/?ref=bushletter.com)
3. [Originality.AI, AI Overview citations vs Google organic rankings](https://originality.ai/blog/google-ranking-ai-citations-study?ref=bushletter.com)
4. [Cloudflare, crawl-to-refer ratios for AI search platforms](https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/?ref=bushletter.com)

FREQUENTLY ASKED QUESTIONS

How many sources does an AI search engine cite per response on average?

The study found LLM-based search engines averaged 4.3 distinct URLs and 3.4 distinct domains per response, compared with 10.3 URLs and 7.3 domains for traditional search engines like Google and Bing.

Why do AI search engines cite so many domains that Google does not rank?

AI engines use retrieval-augmented generation, selecting documents that are linguistically useful for constructing an answer rather than documents that score highly on authority signals like backlinks and click-through rates. Niche publications, specialist databases and forums can appear in AI citations without ever ranking in Google.

Does Grok cite sources in its answers?

Rarely. The study found Grok omitted external citations entirely in 82% of its responses. Gemini did so in 38% of responses. Perplexity was closer to the more citation-heavy end of the LLM range.

What does this mean for SEO strategy?

Because the AI citation pool is small and diverges from Google rankings, third-party coverage matters more than it does in traditional SEO. What independent review sites, publications and forums say about a business may be the primary signal determining whether an AI engine names that business in its answers.

![Takeshi Mori](https://res.cloudinary.com/dz77sb7j1/image/upload/v1774262608/bushletter/authors/takeshi-mori.png)

[Takeshi Mori](https://bushletter.com/author/takeshi-mori/?ref=bushletter.com)

Takeshi Mori writes about technology and start-ups. He is curious about how products get built and who they are really for, and he would rather see a thing working than hear it described.