Subscribe
Marketing

Only about 10% of AI-cited sites also rank top on Google

Researchers ran 1,000 ranking-style queries and measured how often the domains appearing in Google's top-10 results also turned up in citations generated by four major AI platforms.

8 min read
A Google search executive on stage in front of a giant screen reading AI Mode.
Google's AI Mode leans on sites that mostly do not top the classic rankings. | Digitally illustrated image
Elias Thorne
By Elias Thorne · 2026-08-04

TLDR

Ranking at the top of Google and appearing in an AI-generated answer are now largely separate outcomes. Domain-level overlap between Google's top-10 results and AI citations averaged just 10.7% across four major platforms, with GPT-4o's median overlap sitting at zero.

KEY TAKEAWAYS

01Domain overlap between Google's top-10 and AI citations averaged 10.7% across GPT-4o, Gemini, Claude and Perplexity.
02GPT-4o's median overlap with Google rankings was 0%, meaning most queries shared no domains at all.
03Perplexity cited the most sources per prompt at 16.35, but ChatGPT's cited pages carried far higher influence scores.
0486% of AI-cited sources across 6.8 million citations were brand-managed properties, not editorial pages.
05ChatGPT Search rewrites user prompts into multiple targeted sub-queries, pulling sources from across the web.

What the studies actually measured

Researchers ran 1,000 ranking-style queries and measured how often the domains appearing in Google's top-10 results also turned up in citations generated by four major AI platforms. The mean domain-level overlap was 4.0% for GPT-4o, 11.1% for Gemini, 12.6% for Claude and 15.2% for Perplexity, averaging roughly 10.7% across the four systems.[1] By any reasonable measure, the two visibility systems are running almost independently of each other.

GPT-4o produces the sharpest data point. The median domain-level overlap between GPT-4o citations and Google's top-10 results was 0.0%, meaning more than half of all tested queries returned no shared domains whatsoever between the two systems.[1] A brand that has spent years optimising its way to Google's first page cannot assume that work carries over into an AI-generated answer.

Why AI assistants don't start where Google does

The divergence is not accidental. OpenAI's own documentation confirms that ChatGPT Search rewrites user prompts into one or more distinct targeted queries, allowing a single question to be decomposed into multiple sub-queries that each retrieve a different slice of the web.[3] A user asking a broad question about electric vehicle charging infrastructure may trigger separate searches for grid capacity data, manufacturer specifications and policy timelines at the same time.

That multi-step retrieval changes which pages get surfaced. A highly ranked generalist page, the kind built to satisfy a wide keyword, may score well against the original broad prompt but lose out to a narrower, lower-ranked page that precisely answers one sub-query.[1] Google orders pages; AI assistants assemble answers, and those are different problems.

The research is consistent on this point. Generative search exposes a two-stage process: citation selection, where the platform triggers web searches and assembles a pool of candidate sources, and citation absorption, where a selected page meaningfully contributes language, facts or structure to the final answer.[2] Ranking well in Google addresses only the first stage, and only partially at that.

Citation breadth versus citation depth

A separate study of controlled prompts across ChatGPT, Google's AI Overview and Perplexity examined not just how many sources each platform cited but how much those sources actually shaped the answer. Perplexity cited the most sources per prompt, averaging 16.35, while Google's AI Overview averaged 12.06 and ChatGPT returned 6.88.[2] Citing more sources is not the same as being a better optimisation target.

ChatGPT's cited pages carried an average influence score of 0.2713, compared with 0.0584 for Google's AI Overview and 0.0646 for Perplexity, despite ChatGPT citing fewer sources overall.[2] A ChatGPT citation tends to mean the page contributed substantially to the answer text, not merely that it was listed as a reference. Volume and influence move in opposite directions across these platforms.

That distinction matters for anyone trying to measure AI visibility. A brand appearing in 16 Perplexity citations on a single query is not automatically better positioned than one appearing in four ChatGPT citations, given how divergent the influence-per-citation figures are. Counting mentions without weighting them by their contribution to the answer understates how concentrated the real visibility gains are.[2]

Where traditional SEO investment still connects

The overlap is low, but it is not zero, and the composition of AI-cited sources points to where existing digital investment still pays off. A Yext study of 6.8 million AI-generated citations across ChatGPT, Gemini and Perplexity found that 86% of cited sources were brand-managed properties, a company's own website, directory listings or review pages, rather than editorial or third-party pages.[4] Structured, accurate first-party data remains a reliable input into AI retrieval even when organic search ranking is not.

Christian J. Ward, Chief Data Officer at Yext, said discussions about measuring AI visibility are missing the most important factor: the consumer. Ward said AI generates answers based on a person's real-world location and context, not a generic brand view, and that this has led to more confusion than clarity about what really powers AI.[4] That framing shifts the problem away from content optimisation toward data accuracy, specifically whether a business's name, address, product information and structured attributes are consistent and machine-readable across the sources AI platforms draw from.

Mike Walrath, CEO of Yext, said structured, accurate data is the foundation of digital visibility, and that the research confirms what the company has always known: when brands control their data, they control their visibility.[4] The Yext study is vendor-commissioned and its conclusions align with the company's commercial offering, which should be weighed accordingly. The underlying citation-volume figure, 6.8 million citations across three major platforms, does represent one of the larger empirical datasets available on the question.

The measurement gap this creates

Traditional SEO performance is tracked through ranking positions and click-through rates, both of which assume the user navigated from a search results page to a destination URL.[1] Generative AI search routinely satisfies user intent without producing a click at all, synthesising an answer from multiple sources and presenting it directly. A page can contribute materially to an AI answer and generate zero attributable traffic.

The divergence between the two systems is not uniform across query types. The academic research on 1,000 ranking-style queries found that Perplexity maintained the highest overlap with Google's top 10 at 15.2%, suggesting that some query categories, those where authoritative, well-ranked pages are the best factual source, do produce meaningful convergence.[1] The near-zero GPT-4o figures reflect a retrieval approach that appears to weight freshness, specificity and structured data over domain authority.

For marketing and search teams, the data supports treating AI citation and Google ranking as two distinct performance metrics requiring separate measurement frameworks. Pages with specific, data-rich, structured content can surface inside AI answers without ever cracking Google's first page for a broad head term.[3] Pages that rank well on domain authority and backlink profiles alone may find that currency does not carry over into a retrieval environment that decomposes questions before it fetches answers.

FREQUENTLY ASKED QUESTIONS

What does domain-level overlap mean in this context?
It measures how often the same web domains appear in both Google's top-10 organic results and the sources an AI platform cites when answering the same query. A 10.7% average overlap means that roughly 9 in 10 domains cited by AI assistants are not present in Google's top results for the same question.
Why does GPT-4o have a 0% median overlap with Google?
Researchers found that for more than half of the 1,000 test queries, GPT-4o's cited domains shared nothing with Google's top-10 results. OpenAI's own documentation confirms ChatGPT Search breaks queries into multiple targeted sub-queries, which retrieves a wider and more varied source pool than a standard ranked results page.
Does being cited by AI mean your page actually influenced the answer?
Not necessarily. Research comparing citation breadth and influence scores found that Perplexity cites the most sources per prompt (16.35 on average) but those pages carry a far lower influence score than ChatGPT's fewer citations. A citation from ChatGPT is more likely to reflect genuine textual contribution to the answer.
What type of content do AI platforms cite most often?
A study of 6.8 million citations across ChatGPT, Gemini and Perplexity found 86% were brand-managed properties, including company websites, listings and review pages, rather than editorial or independent third-party sources. That figure comes from Yext, a vendor with a commercial interest in structured business data.
Elias Thorne

Elias Thorne

Elias Thorne writes about interest rates, the bond market and the Reserve Bank. He is interested in what monetary policy actually does to household budgets, and in the long stretches of economic history that tend to repeat.

Related topics
What's your reaction?

Make us a preferred source on Google

Tap once and our reporting shows at the top of your Google search results and AI answers. You can change this at any time.

Add as a preferred source on Google
Subscribe — it's free