
TLDR
Ranking at the top of Google and appearing in an AI-generated answer are now largely separate outcomes. Domain-level overlap between Google's top-10 results and AI citations averaged just 10.7% across four major platforms, with GPT-4o's median overlap sitting at zero.
KEY TAKEAWAYS
What the studies actually measured
Researchers ran 1,000 ranking-style queries and measured how often the domains appearing in Google's top-10 results also turned up in citations generated by four major AI platforms. The mean domain-level overlap was 4.0% for GPT-4o, 11.1% for Gemini, 12.6% for Claude and 15.2% for Perplexity, averaging roughly 10.7% across the four systems.[1] By any reasonable measure, the two visibility systems are running almost independently of each other.
GPT-4o produces the sharpest data point. The median domain-level overlap between GPT-4o citations and Google's top-10 results was 0.0%, meaning more than half of all tested queries returned no shared domains whatsoever between the two systems.[1] A brand that has spent years optimising its way to Google's first page cannot assume that work carries over into an AI-generated answer.
Why AI assistants don't start where Google does
The divergence is not accidental. OpenAI's own documentation confirms that ChatGPT Search rewrites user prompts into one or more distinct targeted queries, allowing a single question to be decomposed into multiple sub-queries that each retrieve a different slice of the web.[3] A user asking a broad question about electric vehicle charging infrastructure may trigger separate searches for grid capacity data, manufacturer specifications and policy timelines at the same time.
That multi-step retrieval changes which pages get surfaced. A highly ranked generalist page, the kind built to satisfy a wide keyword, may score well against the original broad prompt but lose out to a narrower, lower-ranked page that precisely answers one sub-query.[1] Google orders pages; AI assistants assemble answers, and those are different problems.
The research is consistent on this point. Generative search exposes a two-stage process: citation selection, where the platform triggers web searches and assembles a pool of candidate sources, and citation absorption, where a selected page meaningfully contributes language, facts or structure to the final answer.[2] Ranking well in Google addresses only the first stage, and only partially at that.
Citation breadth versus citation depth
A separate study of controlled prompts across ChatGPT, Google's AI Overview and Perplexity examined not just how many sources each platform cited but how much those sources actually shaped the answer. Perplexity cited the most sources per prompt, averaging 16.35, while Google's AI Overview averaged 12.06 and ChatGPT returned 6.88.[2] Citing more sources is not the same as being a better optimisation target.
ChatGPT's cited pages carried an average influence score of 0.2713, compared with 0.0584 for Google's AI Overview and 0.0646 for Perplexity, despite ChatGPT citing fewer sources overall.[2] A ChatGPT citation tends to mean the page contributed substantially to the answer text, not merely that it was listed as a reference. Volume and influence move in opposite directions across these platforms.
That distinction matters for anyone trying to measure AI visibility. A brand appearing in 16 Perplexity citations on a single query is not automatically better positioned than one appearing in four ChatGPT citations, given how divergent the influence-per-citation figures are. Counting mentions without weighting them by their contribution to the answer understates how concentrated the real visibility gains are.[2]
Where traditional SEO investment still connects
The overlap is low, but it is not zero, and the composition of AI-cited sources points to where existing digital investment still pays off. A Yext study of 6.8 million AI-generated citations across ChatGPT, Gemini and Perplexity found that 86% of cited sources were brand-managed properties, a company's own website, directory listings or review pages, rather than editorial or third-party pages.[4] Structured, accurate first-party data remains a reliable input into AI retrieval even when organic search ranking is not.
Christian J. Ward, Chief Data Officer at Yext, said discussions about measuring AI visibility are missing the most important factor: the consumer. Ward said AI generates answers based on a person's real-world location and context, not a generic brand view, and that this has led to more confusion than clarity about what really powers AI.[4] That framing shifts the problem away from content optimisation toward data accuracy, specifically whether a business's name, address, product information and structured attributes are consistent and machine-readable across the sources AI platforms draw from.
Mike Walrath, CEO of Yext, said structured, accurate data is the foundation of digital visibility, and that the research confirms what the company has always known: when brands control their data, they control their visibility.[4] The Yext study is vendor-commissioned and its conclusions align with the company's commercial offering, which should be weighed accordingly. The underlying citation-volume figure, 6.8 million citations across three major platforms, does represent one of the larger empirical datasets available on the question.
The measurement gap this creates
Traditional SEO performance is tracked through ranking positions and click-through rates, both of which assume the user navigated from a search results page to a destination URL.[1] Generative AI search routinely satisfies user intent without producing a click at all, synthesising an answer from multiple sources and presenting it directly. A page can contribute materially to an AI answer and generate zero attributable traffic.
The divergence between the two systems is not uniform across query types. The academic research on 1,000 ranking-style queries found that Perplexity maintained the highest overlap with Google's top 10 at 15.2%, suggesting that some query categories, those where authoritative, well-ranked pages are the best factual source, do produce meaningful convergence.[1] The near-zero GPT-4o figures reflect a retrieval approach that appears to weight freshness, specificity and structured data over domain authority.
For marketing and search teams, the data supports treating AI citation and Google ranking as two distinct performance metrics requiring separate measurement frameworks. Pages with specific, data-rich, structured content can surface inside AI answers without ever cracking Google's first page for a broad head term.[3] Pages that rank well on domain authority and backlink profiles alone may find that currency does not carry over into a retrieval environment that decomposes questions before it fetches answers.
SOURCES & CITATIONS
FREQUENTLY ASKED QUESTIONS
What does domain-level overlap mean in this context?
Why does GPT-4o have a 0% median overlap with Google?
Does being cited by AI mean your page actually influenced the answer?
What type of content do AI platforms cite most often?

Elias Thorne writes about interest rates, the bond market and the Reserve Bank. He is interested in what monetary policy actually does to household budgets, and in the long stretches of economic history that tend to repeat.



