Menu

Menu

Menu

Menu

Jul 15, 2026

GEO vs SEO: What AI Engines Actually Cite [2026 Study]

Jul 15, 2026

GEO vs SEO: What AI Engines Actually Cite [2026 Study]

We asked Claude, GPT, and Gemini 96 real questions, four times each on the same day, through their live, public search tools, and recorded every step of how each answer was built: the searches issued, the sources seen, the sources cited, and the brands named. Every search was then checked against same-day results from Brave, Bing, and Google, collected 100 results deep.

START FREE. NO CREDIT CARD REQUIRED.

START FREE. NO CREDIT CARD REQUIRED.

Turn AI Search into growth channel
Turn AI Search into growth channel

Track how ChatGPT, Alexa for shopping, Gemini, Perplexity, Claude, and AI Overviews mention and recommend your brand, strategize your next moves with a GEO agent and our experts, then publish content AI cites.

Track how ChatGPT, Alexa for shopping, Gemini, Perplexity, Claude, and AI Overviews mention and recommend your brand, strategize your next moves with a GEO agent and our experts, then publish content AI cites.

What AI answer engines recommend and cite

We asked Claude, GPT, and Gemini 96 real questions, four times each on the same day, through their live, public search tools, and recorded every step of how each answer was built: the searches issued, the sources seen, the sources cited, and the brands named. Every search was then checked against same-day results from Brave, Bing, and Google, collected 100 results deep.

The central question: how much of an AI engine's citation behavior does classic search ranking still explain, and how much is a genuinely new game? That breaks into the questions this report measures one by one:

  • Ranking. How much of what gets cited was already ranking (SDS), and how much ranks nowhere on any of the three search engines (GEO Surplus).

  • Backends. Which search engine actually explains each AI's citations, tested against all three rather than assumed, and whether two AIs disagree because of their search engines or their own judgment.

  • Selection. How much of what an engine sees it leaves uncited, and how sharply the odds of citation fall with search position.

  • Rewriting. How far engines rewrite the question before searching, and how much that rewrite changes which sources win the citation.

  • Stability. How much answers change when the identical question is asked again, and how every pattern above bends by question type.

  • Sources. What kind of evidence each engine trusts: editorial, vendor pages, forums and video; how concentrated citations are; and whether cited domains are ones search engines recognize at all.

  • Brands. Which brands get recommended by name against which domains get cited as evidence. Those turn out to be two different wins, needing two different playbooks.

Key numbers: 96 prompts · 1,152 answers · 3 engines tested · 8,278 entities extracted · 24,931 citations


01 · Definitions

Ten terms used throughout this report

Ten terms recur throughout this report. Each one measures something specific about how an AI engine decides what to cite and recommend. Read this table once and the rest of the report needs no further explanation of what a number means, only what it shows.

The ten terms, at a glance


Term

What it measures

SDS (SERP Dependence Score)

The share of an engine's citations that already rank within a given depth of real search results, on the search engine that best explains its citations.

GEO Surplus

The share of citations that rank nowhere, on any of the three search engines checked, at 100 results deep. The part search ranking cannot explain.

Selection Gap

The share of available sources an engine never cites. Measured against its own considered list where visible, and against real search results everywhere else.

Query Drift

How far an engine's actual search terms move from the user's original question, and how much that movement changes which sources get credited with a citation.

Answer Churn

How much an engine's citations change when the identical question is asked again on the same day.

Index Attribution

Which of the three search engines, Brave, Bing, or Google, best explains a given AI engine's citations, tested directly rather than assumed.

Domain Class

The category a cited source falls into: editorial, brand or vendor, official, forum or social, reference, or documentation.

Brand Self-Citation Rate

How often a brand recommended by name also has its own website among the cited sources.

Citation Concentration

How much of all citations are captured by just the ten most-cited domains.

Source Recognition

Whether a cited domain shows up anywhere in the relevant search engine's own results, used as a stand-in for site authority since backlink data was not available.


02 · Methodology

What we ran, and what we measured

This is what we ran, on what, and what we measured. Everything else in this report follows from these choices.

We sent 96 English-language prompts across eight question types (broad commercial, specific commercial, brand comparisons, fresh news, stable facts, branded lookups, local search, and contested topics) to Claude, GPT, and Gemini. Each prompt was asked four times on the same day to each of the three engines: 12 answers per prompt, 1,152 in total. Each engine used its own live, public search tool: no shared backend, no simulated results.

Every answer went through the same three steps, and we recorded all three: the engine first decided whether to search at all, then issued zero or more of its own search queries, then wrote an answer citing some subset of what it found. Claude (claude-sonnet-4-6) and GPT both expose the list of sources they considered before writing, so we captured that candidate list alongside the queries and the final citations; Gemini (gemini-3.5-flash) does not expose a pre-citation candidate list, only the queries and the final citations, which is why this report uses a SERP-inferred proxy rather than a native selection-gap number for Gemini.

For every query any engine issued, regardless of which engine issued it, we separately queried all three search indexes ourselves, down to 100 results deep, on the same day: Brave direct through its own search API, Bing through Bright Data, and Google through DataForSEO's Live Advanced product (DataForSEO's own Bing product returned unusable, wrong-locale results in testing, which is why Bing runs through a different provider than Google). That gives three independent, same-day snapshots of what actually ranked for every query, on every index, regardless of which search engine an engine is assumed to run on. Every claim in this report about ranking, SDS, GEO Surplus, and the index-attribution matrix included, is a lookup against these three snapshots, not an assumption about what any engine's backend returned.

Separately, we ran the full text of every answer through an LLM extraction pass to pull out every brand, product, and company name mentioned, whether or not that brand's own site was ever cited. That is the raw material behind the recommend-vs-cite comparison later in this report.

What this does not cover: the consumer apps (ChatGPT.com, the Gemini app, Claude.ai), which may behave differently from the developer tools tested here; languages other than English; and behavior beyond a single day. It also shows what gets cited relative to what ranks, not why one page is chosen over another.

What each engine ran


Engine

Model

Search tool

Best-matching search engine

Tested against

Note

Claude

Claude Sonnet 4.6 (default, free-tier)

Native web search

Brave

Brave, Bing, Google

assumed and measured match

GPT

GPT-5.5 Instant (default, free-tier)

Native web search

Brave

Brave, Bing, Google

assumed Bing, measured Brave

Gemini

Gemini 3.5 Flash (default, free-tier)

Native search grounding

Google

Brave, Bing, Google

assumed and measured match

Where the search-ranking data came from. For every query any engine issued, we queried all three indexes below ourselves, independent of which engine issued the query or which index it's assumed to run on.

Search-index collection, per index


Search index

Provider

Results depth

Note

Brave

Brave Search API (direct)

100

Claude's matched index

Bing

Bright Data

100

GPT's assumed index; DataForSEO's Bing product returned unusable results in testing

Google

DataForSEO Live Advanced

100

Gemini's matched index

We successfully collected usable search results for Brave 100%, Bing 99%, Google 98% of searches. The small remainder failed to return enough results to use and was excluded rather than guessed at.


03 · Key Results

Search rank explains almost all of Claude's citations, and barely two-thirds of GPT's

Search ranking explains what gets cited to a different degree for every engine we tested, and the gap between the extremes is larger than any engine's own margin of error.

Claude's SDS at 100 results deep is 96.5 percent, with a GEO Surplus of 1.6 percent. GPT's SDS is 63.9 percent, with a GEO Surplus of 20.7 percent. Gemini sits between the two, at 89.2 percent and 5.0 percent. The Claude-to-GPT gap holds up after accounting for random variation between prompts (a 43-point difference in SDS at the top ten results, likely between 37 and 50 points), so this is not sampling noise.

Volume and search-ranking dependence move in opposite directions. Claude and Gemini cite heavily (6.5 and 10.4 unique domains per answer) and lean hard on search rankings. GPT cites sparingly (2.6 domains per answer) and leans on rankings the least. That restraint is the mechanism behind its GEO Surplus: it searches on 68 percent of prompts, at 0.7 searches per answer on average, and the small set it ends up citing reaches beyond the rankings more often than either rival's much larger set.

Index Attribution tells a related story from a different angle. Claude lines up clearly with Brave and Gemini with Google, matching the common assumption for both. GPT does not cleanly match any of the three: its citations line up with Brave, Google, and Bing within five points of each other at every depth, and Bing, the search engine it is commonly assumed to run on, explains it worst of the three. Ranking first on Google reaches only one of the three engines tested here.

In one line each: Claude searches on 87.5 percent of prompts and cites an already-ranking, editorial-and-vendor-heavy set with the lowest day-to-day variation of the three (24 percent). GPT searches and cites the least, and reaches beyond the rankings the most, leaning on Reddit and Wikipedia when it does. Gemini searches almost every time a question needs current information, issues the most searches per prompt at nearly five, and cites more sources than either rival by a wide margin (46 per answer on average), largely through Reddit and YouTube.

How much of an AI engine's citation behavior classic search ranking explains has no single answer. For Claude it is nearly all of it. For GPT it is roughly two-thirds, with a real, sizable remainder no ranking strategy can reach. A domain that ranks on Google but not Brave should show up far less often in Claude's citations than in Gemini's. GPT is the one engine of the three where content and source selection matter more than rank, since a fifth of what it cites was never going to rank in the first place.

Chart: SDS vs GEO Surplus at 100 results deep, per engine

04 · Recommend vs Cite

Most recommended brands win without their own site ever being cited

Being recommended by name and being cited as a source are two separate wins, and most studies only measure the second one. When an answer says "the best option is X," that is a recommendation. When it links to a page as evidence, that is a citation. The brands an engine recommends are frequently not the domains it links to.

The Brand Self-Citation Rate captures this directly: of the brands an engine recommends by name, how often does that brand's own website appear among the cited sources? Counting each recommended brand once per answer, Claude sits at 10 percent, GPT at 14 percent, Gemini at 22 percent. Even on Gemini, the highest of the three, roughly four out of five recommended brands are endorsed without their own site ever being cited. The recommendation is won through someone else's coverage: reviews, forums, comparisons, not the brand's own page.

Engines also name far more brands than they ever cite. Gemini mentions about nine per answer, Claude eight, GPT five, and most of those are actual brand or product names rather than generic terms. A brand can win the recommendation without appearing in the source list at all.

Recommendation rewards brand recognition the model already has for the topic at hand. Citation rewards third-party evidence. They are different levers, and building one does not automatically build the other.

A mention is an entity the answer names as something to choose (e.g. "Adidas Adizero"). A citation is a source domain linked as evidence (e.g. runrepeat.com). A brand wins by being mentioned; a publisher wins by being cited.

Chart: Brand Self-Citation Rate


Implication for GEO Winning AI recommendations means influencing the sources an engine reads: review sites, forums, and video, not only a brand's own domain. Ranking a brand's own site moves the citation game far more than the recommendation game.


05 · Source Mix

There is no single AI-friendly source type: each engine trusts different evidence

The Domain Class of what gets cited is its own fingerprint, and the three engines look nothing alike. We sorted every cited domain into editorial, brand or vendor, official, forum or social, reference, or documentation, and measured both the share of citations each class makes up and the share of answers that include at least one.

Claude is editorial-led: editorial sites make up 43 percent of its citations, brand pages 39 percent, and everything else trails far behind. GPT leans similarly on editorial and brand sources but draws more from official and technical documentation than the other two, and cites a forum or social source in only 13 percent of its answers. Gemini is different in kind, not just degree: brand pages lead at 39 percent, editorial follows at 26 percent, and forums or social posts make up 23 percent, four times Claude's share. By answer, 74 percent of Gemini's responses cite at least one forum or social post, against 32 percent for Claude and 13 percent for GPT.

The individual publishers make the pattern concrete. Reddit appears in 50 percent of Gemini's answers and YouTube in 49 percent; between them they account for nearly all of Gemini's forum and social citations. GPT cites Reddit too, but at 11 percent, alongside TechRadar, Wikipedia, and Tom's Guide as one source among several rather than a dominant pair. Claude cites none of them: its top sources are AOL, Tom's Guide, Capterra, and TechRadar, tech-review sites and software directories with no forum or video presence at all.

There is no single AI-friendly source type. Winning on Gemini means earning Reddit threads and YouTube coverage. Winning on Claude and GPT means tech-review sites, software directories, and product pages. A content strategy built for one of these engines will land poorly on another.

Citation Concentration adds a second dimension to the same picture: how much of an engine's citations come from just its ten most-cited domains, compared to how concentrated the underlying search results are. Claude's citations are about as concentrated as its search results (8 percent versus 9 percent), meaning it draws from roughly as wide a pool as the rankings offer. GPT and Gemini both cite more narrowly than their search results would suggest, at 23 percent and 23 percent respectively against 10 percent and 17 percent for the rankings themselves. Both engines have settled on a smaller set of trusted sources than the breadth of what actually ranks for their searches.

Every cited domain is classified into one Domain Class. The engines ride different search engines, and they also prioritize results differently, so the evidence mix diverges even on identical prompts.

Chart: Citation mix by source type (share of each engine's cited instances by source kind, normalized to 100 percent)


Chart: Reach of forums, video, and Q&A (share of answers that cite at least one forum, video, or Q&A source)


Claude: top cited domains


Domain

% answers

aol.com

9%

tomsguide.com

7%

capterra.com

5%

techradar.com

4%

tech-insider.org

4%

bankrate.com

3%

GPT: top cited domains


Domain

% answers

reddit.com

11%

techradar.com

10%

en.wikipedia.org

6%

tomsguide.com

6%

forbes.com

5%

axios.com

3%

Gemini: top cited domains


Domain

% answers

reddit.com

50%

youtube.com

49%

medium.com

10%

en.wikipedia.org

9%

forbes.com

8%

cnet.com

5%

Citation Concentration: share held by the ten most-cited domains


Engine

Share of citations

Share of search results

Claude

8%

9%

GPT

23%

10%

Gemini

23%

17%


06 · Retrieval

Engines answer their own rewritten searches, not the question you typed

Before an engine runs a search, it decides what to search for, and that is rarely the words a person typed. Optimizing for someone's literal question misses the target the engine actually searches.

The three engines research in distinct styles. Gemini digs hardest, at nearly five searches per prompt and up to fifteen on some questions, treating a single question like a small research project. Claude is moderate, at about two searches per prompt, and follows up: it runs a search, reads the results, and often searches again based on what it found. GPT is minimal, at roughly one search per prompt and often none at all, averaging under one search per answer.

Query Drift measures how much the actual search terms move from the original question. On average, roughly half to two-thirds of the wording changes between what a person asked and what gets searched, with GPT rewriting the most aggressively despite searching the least. That movement has a measurable consequence: search rank predicts citation far better against the engine's actual search terms than against the original question. For Claude, 81 percent of citations line up with a top-ten result when measured against its own search terms, against 42 percent when measured against the original question as if it were the search itself. For Gemini the gap is 63 percent versus 34 percent, and for GPT 44 percent versus 30 percent, smaller mainly because GPT's overall connection to search rank is weaker to begin with. The searches that matter are the rewritten ones, not the question a person typed, and those rewritten searches are invisible from the prompt alone.

For the two engines that expose their full result list before selection, citations come from across that list, not only the top. Claude sees about thirteen results per answer and cites fairly evenly from the top three (26 percent of citations), the middle (31 percent from positions six through ten), and further down (28 percent from position eleven or later); it still cites from deep in its own list four times out of ten. GPT sees more results, about sixteen, but its citation rate drops sharply with position: it cites half of what its results place in the top three, and under one in ten of what lands at position eleven or later. It looks deep and cites shallow. Gemini exposes only the sources it ended up using, not its full result list, so this comparison is not possible for Gemini.

The cleanly comparable signal is query fan-out: how many searches an engine issues, split by whether the prompt has fresh, time-sensitive intent or stable intent.

Chart: Queries issued per answer (mean queries issued per answer, by intent class; Gemini fans out hardest, Claude is the most economical)


Chart: Where cited sources sat in the engine's own results (normalized; Claude and GPT only, Gemini only shows the sources it ended up using, not its full result list)


Chart: Cite rate by position in the engine's own results (of the sources an engine's own results surface at each rank, the share it ends up citing; GPT drops off steeply as rank falls, Claude keeps citing further down its own short list)


Search terms versus the original question

SDS measured against an engine's real search terms compared to SDS measured against the original question as if it were the search itself. A bigger gap means the rewritten search explains far more than the words a person typed.


07 · Search Reliance

Rank is a strong nudge toward citation, but never a gate

SDS measured at increasing depth shows how much of an engine's citation behavior collapses into search ranking once you are willing to look further down the page. Claude's SDS runs 80 percent at the top ten, 93 percent at the top twenty, and 96.5 percent at the top hundred: already near-complete at the first page, with almost nothing left to explain by looking deeper. Gemini runs 63, 78, and 89 percent: strongly tied to rank, with a real slice of citations that only surface once you look well past the first page. GPT runs 37, 46, and 64 percent, the flattest curve of the three. Even a hundred results deep, more than a third of what GPT cites never appears in the rankings at all.

Each of those headline numbers is measured against the one search engine that fits best. Checked against the other two, Claude's citations sit well behind Brave at every depth (41, 55, and 74 percent against Google; 35, 42, and 65 percent against Bing), and so do Gemini's against anything but Google (54, 65, and 81 percent against Brave; 37, 45, and 68 percent against Bing). GPT is the exception: its numbers against Brave, Google, and Bing sit within five points of each other at every depth, instead of separating the way Claude's and Gemini's do. Most of that overlap is shared, generic content rather than a real tie to Brave. Breaking down why GPT's citations rank on Brave: at the top ten, only about six of the thirty-seven points come from pages that rank on Brave and nowhere else; the other thirty-one also rank on Bing or Google, so they would have looked like a match against any of the three. Claude's equivalent split at the top ten is twenty-nine Brave-exclusive points out of eighty, roughly five times GPT's exclusive share. Close to half of GPT's citations, at the top ten, rank on none of the three search engines checked, which argues against a clean story of GPT blending exactly these three sources. What is actually behind that pattern, a blended index, a private crawl, or something else entirely, cannot be confirmed from citation overlap alone.

Rank still predicts citation, just weakly, and never completely. Claude's citation rate declines from 53 percent at position one through three to 11 percent at position fifty-one through a hundred, the steadiest drop of the three. Gemini declines from 58 to 36 to 24 percent, with an unusual rebound to 24 percent again at the very bottom of the range, meaning its deep results get cited about as often as its mid-ranked ones. GPT declines from 30 to 12 percent, the flattest of the three; rank barely moves its citation odds at all. A top-three ranking roughly doubles the odds of being cited compared to the deep end of the results, but even a first-place ranking is cited well under 60 percent of the time, and even the deepest results still get cited 10 to 24 percent of the time. Rank is a strong nudge toward citation. It is not a gate.

Compared against real search results from Brave, Bing, and Google, collected up to 100 results deep, for each engine's own searches. Exact page means the specific cited page shows up in those results. Whole site means any page on that same domain counts. Shown at several depths, from the top 10 through the top 100, so both the exact-page-versus-whole-site distinction and the effect of looking deeper are visible, rather than squeezed into one number.

Chart: SDS by depth, top 10 through top 100


SDS at each depth: exact page vs whole site


Engine

Search engine

Match type

≤10

≤20

≤30

≤50

≤100

Searches covered

Claude

Brave

Exact page

76%

90%

91%

93%

93%

100%

Claude

Brave

Whole site

80%

93%

94%

96%

96%

100%

GPT

Brave

Exact page

20%

24%

27%

32%

36%

100%

GPT

Brave

Whole site

37%

46%

52%

58%

64%

100%

Gemini

Google

Exact page

45%

58%

63%

68%

71%

99%

Gemini

Google

Whole site

62%

78%

84%

88%

89%

99%

The same measurement, against every search engine, not just the best match

One engine per card: how well its citations line up with Brave, Bing, and Google individually, at each depth. Lines that stay close together, as with GPT, mean the best-match score above is mostly shared, generic content rather than a real preference for one search engine.

High lines against several search engines at once do not contradict the low overlap between the search engines themselves (they share under a third of their results, even 100 deep). That overlap counts every ranked domain equally, long tail included, while citations concentrate in a small head of popular domains that rank on all three, and each answer's check pools every search it ran. Gemini scoring 89% against Google and 81% against Brave means what it cites is mostly content that ranks everywhere; the true backend only shows in the slices that rank on one search engine and not another, which is exactly what the attribution section below isolates.

Claude's citations, checked against each of the three search engines at every depth

Search engine

≤10

≤20

≤30

≤50

≤100

Brave ◆ best match

80%

93%

94%

96%

96%

Bing

35%

42%

45%

53%

65%

Google

41%

55%

63%

71%

74%


GPT's citations, checked against each of the three search engines at every depth

Search engine

≤10

≤20

≤30

≤50

≤100

Brave ◆ best match

37%

46%

52%

58%

64%

Bing

32%

39%

43%

49%

58%

Google

34%

44%

49%

55%

59%


Gemini's citations, checked against each of the three search engines at every depth

Search engine

≤10

≤20

≤30

≤50

≤100

Brave

54%

65%

70%

76%

81%

Bing

37%

45%

51%

58%

68%

Google ◆ best match

62%

78%

84%

88%

89%

Chart: Chance of citation by search-result position (for each position bucket, the share of ranking domains the engine ends up citing; because results were collected 100 deep, even the deepest bucket is measured directly rather than extrapolated)


08 · What Gets Left Out

Ranking gets a source considered, but most of what engines see goes uncited

Of everything an engine could have cited, how much did it actually leave out? Ranking well gets a page into consideration. A separate step then decides what survives into the answer, and the Selection Gap measures how large that step is.

We measure it two ways, since the engines do not expose the same information. For Claude and GPT, we can see the full list of sources each one considered before writing, so we measure directly how much of that list went unused. For Gemini, that list is not visible, only the sources it ended up using, so we reconstruct the likely pool from real search results for the same queries and measure how much of that pool went uncited. We apply the second method to all three engines so the numbers stay comparable.

Against their own considered list, Claude uses about 63 percent of what it sees and leaves 37 percent unused. GPT uses only about 30 percent, discarding more than twice Claude's share. Its considered list may already be filtered somewhat before we see it, so 70 percent is a floor on how much it actually turns down, not necessarily the full figure.

Against the reconstructed search-result pool, the pattern holds at a larger scale, and grows with depth as the pool grows faster than the citations do. GPT leaves 83 percent of the top ten unused, rising to 95 percent by the top hundred, the widest gap of the three. Gemini leaves 67 percent of the top ten unused, rising to 92 percent, despite citing more than either rival. Claude leaves 58 percent of the top ten unused, the narrowest gap at that depth, rising to 92 percent by the top hundred.

Every engine leaves out most of what it could cite, and the ordering runs opposite to how closely each one tracks rank: Claude, which depends most on rank, wastes the least of what it sees; GPT, which depends on rank the least, discards the most. Ranking gets a source into consideration. It does not decide what makes it into the answer. On Gemini and GPT especially, appearing in the results is the easy part. What actually decides survival, position, brand recognition, or the content itself, is a harder question this kind of data cannot answer on its own.

One caution on the raw percentages: the share left unused depends both on how large the result pool grows and on how few sources an engine ends up citing, so it is not a clean measure of selectivity by itself. GPT runs the fewest searches per answer of the three yet posts the highest unused rate by the top hundred, because it cites so few sources overall, not because its pool is unusually large. The trend across depth tells you more than any single figure.

Own-list gap is the share of the engine's own list of considered sources that never got cited, for Claude and GPT only, since only they expose that list. Search-pool gap is the same idea for all three engines: how much of everything that shows up in a real search for the same queries goes uncited, at different depths into the results, from the top 10 through the top 100.

Chart: Share of real search results left uncited, by depth (higher means more of the pool at that depth went unused)


Selection Gap: own list vs search pool


Engine

Own list

Search pool @10

@20

@30

@50

@100

Claude

37%

58%

73%

80%

87%

92%

GPT

70%

83%

89%

91%

93%

95%

Gemini

n/a

67%

79%

84%

89%

92%

Search-pool gap at the top 10 results: likely range Claude: 58% (56%–61%) · GPT: 83% (81%–86%) · Gemini: 67% (65%–69%)


09 · GEO Surplus

A fifth of what GPT cites ranks nowhere at all; on Claude that surface barely exists

GEO Surplus measures how much of what gets cited does not rank at all, not on Brave, not on Bing, not on Google, at a hundred results deep on every search the engine ran. It is the part of winning AI citation that no amount of search-ranking effort can buy, because it was never ranking to begin with.

GPT's GEO Surplus is 20.7 percent: one cited domain in five does not rank anywhere we checked. This is the only engine of the three where citing off-ranking sources is a common pattern rather than a rare exception. Gemini's is 5.0 percent, a real but small residual; rank still explains nineteen of every twenty domains it cites. Claude's is 1.6 percent, close to zero: for Claude, ranking somewhere is nearly a precondition for being cited.

These numbers are not "100 minus SDS," and they are not meant to add up that way. SDS measures ranking on one search engine: whichever one best explains an engine's citations. GEO Surplus measures ranking on none of the three we checked, together. A citation can fail the first test while still passing the second, simply by ranking on a different search engine than the one it is best matched to. The index-attribution section below breaks this split out exactly, at the level of unique cited domains: for Gemini, 89 percent of cited domains rank on Google, matching its SDS exactly, a further 5 percent rank on Brave or Bing but not Google, and the remaining 6 percent rank on none of the three. For GPT, 64 percent rank on Brave, a further 11 percent rank on Bing or Google but not Brave, and the remaining 25 percent rank on none of the three. That last figure, computed per unique domain, sits close to but not exactly on top of the 20.7 percent GEO Surplus figure above, which counts every individual citation rather than each domain once. A domain cited several times in one answer counts once in the domain-level version and several times in the citation-instance version used as this report's headline GEO Surplus number; the two ways of counting agree closely, not exactly.

GEO Surplus is measured only at a hundred results deep on purpose. That depth is the point where "does not rank" cleanly means absent, rather than merely ranking further down the page. The related question, how much of an engine's citation behavior collapses into rank as you look deeper, is answered by SDS elsewhere in this report, not by this metric: SDS climbs from 37 to 64 percent for GPT, 80 to 96.5 percent for Claude, and 63 to 89 percent for Gemini between the top ten and the top hundred.

GPT's non-ranking citations concentrate in commercial questions: brand comparisons (33 percent), specific commercial searches (26 percent), broad commercial searches (26 percent), and local searches (24 percent), and are close to zero on branded lookups or questions with conflicting evidence, where GPT barely searches at all. The sources themselves mix genuine authority with thin content: alongside AP News, TechRadar, Wikipedia, and Forbes sit much smaller sites with little independent standing. We cannot yet tell which of these represent real opportunity and which are simply thin content that happened to match a narrow search, since a working authority score was not available for this run and our source-type classifier sometimes mislabels these same thin sites as editorial or vendor content. Until that gap is closed, GPT's 20.7 percent should be read as citations that do not rank, not as confirmed opportunity.

Search ranking buys the right to be considered on every engine, and the non-ranking surface stays small almost everywhere. On GPT, a fifth of what gets cited is won entirely outside the rankings, exactly where a ranking-only strategy cannot see. On Claude, that surface barely exists.

Already ranking = citations found somewhere in that engine's real search results. GEO Surplus (measured 100 results deep) splits exact page (the specific cited page is absent from every search engine's top 100 results, across every search the engine ran) vs. whole site (no page on that domain shows up at all). Best-matching search engine, measured: Claude→Brave · GPT→Brave · Gemini→Google. Each figure shows our best estimate, with the likely range it falls in.

GEO Surplus with likely ranges, per engine


Engine

GEO Surplus (exact page)

GEO Surplus (whole site)

Already ranking, top 10

Claude

6% (4%–8%)

2% (1%–2%)

80% (77%–83%)

GPT

51% (45%–56%)

21% (16%–26%)

37% (32%–42%)

Gemini

16% (13%–19%)

5% (3%–7%)

62% (59%–66%)


10 · Whose Fault Is the Disagreement?

Different search engines explain most cross-engine disagreement, and the models add more on top

When two AI engines cite very different sources for the same question, that could mean the models themselves disagree, or it could mean they are simply pulling from search engines that already disagree with each other. We separate the two by comparing how much two AI engines actually agree against how much their underlying search engines already agree on the identical searches.

Claude and Gemini agree with each other on about 23 percent of cited sources. Their search engines, Brave and Google, already agree on about 31 percent of results for the same searches. Claude and Gemini disagree more than their search engines would predict, by roughly 8.5 points, which can only come from the models' own choices, not from using different backends. This is the clearest evidence in the study that model judgment, not just search engine choice, drives citations apart.

Claude and GPT agree on about 19 percent of sources, and their search engines, Brave and Bing, already agree on about 20 percent. The two models disagree almost exactly as much as Brave and Bing do: essentially all of their divergence traces back to using different search engines, not to different behavior. GPT and Gemini agree on about 19 percent, while their search engines, Bing and Google, agree on only about 15 percent. Here the models converge slightly more than their search engines do, likely because both lean on the same handful of large, popular sources.

Most of the aggregate disagreement between engines traces back to using different search engines: Brave, Bing, and Google rank different things for the same query. That is not the whole story. Claude and Gemini add real disagreement on top of what their search engines already explain, direct evidence that model-level judgment fragments AI citation beyond what the underlying rankings would predict. Ranking once and winning everywhere fails on two counts here: the search engines disagree, and where they do not, the models sometimes disagree even more.

Even asking the same engine the same question twice on the same day produces different citations, so a small part of any cross-engine gap is ordinary variation rather than a real difference. That variation is far smaller than the 19 to 23 percent gaps measured here, so it does not explain away what we are seeing.

"How much the underlying search engines already overlap" subtracted from "how much these two AI engines actually agree" gives you how much extra disagreement the AI itself adds on top of using different search engines.

Actual vs expected agreement, per engine pair


Engine pair

Actual agreement

Expected from search engines alone

Extra, added by the AI

Claude ∩ GPT

19%

20%

1%

Claude ∩ Gemini

23%

31%

9%

GPT ∩ Gemini

19%

15%

0%


11 · Are These Differences Real?

The gaps between engines are real, not sampling noise

The differences between engines described throughout this report hold up after accounting for random variation between prompts. For SDS at the top ten, every pairwise comparison is a real, distinguishable gap, not a coincidence of which 96 prompts we happened to test.

Claude leads GPT by 43 points, likely between 37 and 50. Claude leads Gemini by 18 points, likely between 14 and 21. GPT trails Gemini by 26 points, likely between 19 and 32. The ordering, Claude ahead of Gemini ahead of GPT, is stable across every comparison, and the likely ranges are tight relative to the size of the gaps: these differences are large enough that they would not disappear with a different sample of prompts.

Each row shows the size of the gap between two engines, plus the range we're confident the true gap falls in.

Pairwise engine gaps: SDS at the top 10


Metric

Engines compared

Gap (likely range)

sds_domain_at_10

claude vs gpt

43% (37%–50%)

sds_domain_at_10

claude vs gemini

18% (14%–21%)

sds_domain_at_10

gpt vs gemini

-26% (-32%–-19%)


12 · Index Attribution

Claude is tied to Brave and Gemini to Google; GPT matches no single search engine

Claude lines up clearly with Brave and Gemini with Google, matching the common assumption for both. GPT does not cleanly match any of the three search engines, with Brave sitting narrowly ahead of Google and Bing, its assumed backend, trailing behind both. The full breakdown by depth is in the search-reliance section above. Here we split each engine's SDS itself into two pieces: of the domains ranking on its best-matching search engine, how many rank there exclusively, and how many also rank on one of the other two.

Claude's SDS at a hundred results is 96.5 percent. That splits into 13 percent that rank on Brave exclusively and 83 percent that rank on Brave and at least one other search engine too (13 plus 83 is the 96.5 already reported). The remaining 3.5 percent rank elsewhere or nowhere. When Claude cites a page that ranks, it is very likely a page that ranks broadly, not a Brave-specific find. Gemini's 89.2 percent splits similarly: 9 percent Google-exclusive, 81 percent Google and elsewhere (9 plus 81 rounds to the 89.2 already reported), and 11 percent combined elsewhere or nowhere.

GPT is the exception, and its split looks nothing like the other two. Of its 63.9 percent SDS, only 5 percent is Brave-exclusive; the other 59 percent also ranks elsewhere (5 plus 59 rounds to the 63.9 already reported). Outside that SDS figure entirely, 11 percent ranks on another search engine but not Brave, and a full 25 percent ranks on none of the three, the same figure reported as GEO Surplus at the domain level in the previous section. That 25 percent is larger than Claude's and Gemini's equivalent figures combined. GPT is the only engine of the three that draws a meaningful share of its citations from clearly outside what any single search engine, or all three together, can explain.

The domains behind that gap are telling. For GPT, sources that rank elsewhere but not on Brave mix recognizable names (AP News, TechRadar, Wikipedia) with much smaller sites. Claude's rare cases in the same bucket skew toward small, niche pages. The contrast holds up the broader pattern: GPT's selection reaches well past any single search engine's results, and Claude's barely does.

The common assumption is Claude with Brave, GPT with Bing, Gemini with Google. Best-matching search engine, measured: Claude→Brave · GPT→Brave · Gemini→Google. An engine pulling from more than one search engine would cite domains that also rank on the other two, not just its best match. We ran every engine's own searches against all three search engines and measured how well each one's citations line up with each. Raw overlap alone can be misleading: if the search engines simply share the same top results for a query, since Wikipedia ranks everywhere, for instance, an engine using only one search engine will still look like it draws from several. The search engines themselves already overlap somewhat on the same searches (comparing their top-10 results, Brave and Bing share 20%, Brave and Google share 31%, and Bing and Google share 15%), so some of what looks like an engine pulling from multiple search engines is really just popular content that ranks everywhere. That is exactly why the chart below isolates the ranks-elsewhere-but-not-here slice. That overlap stays essentially flat all the way to 100 results deep (19%, 28%, and 15% respectively), so the three search engines are not sharing a head and diverging in the tail: they disagree just as much at position 100 as at position 10, and looking deeper buys no extra agreement between them.

Chart: Where cited domains rank across the three search engines (top 10)

The "on another search engine only" slice is the only signal shared content cannot produce on its own. The "best match and others too" slice is the part raw coupling over-counts.

What this test shows

  • Claude is cleanly tied to Brave. Citations that also show up on other search engines almost always show up on Brave too, at 51%; only 4% rank somewhere else while missing from Brave.

  • GPT genuinely pulls from more than one search engine. 16% of its cited domains rank on another search engine but not on Brave, versus only 6% that rank on Brave alone, a pattern shared, popular content cannot explain away, and 48% rank in none of the three search engines' top 10 at all.

  • Gemini is cleanly tied to Google. Citations that also show up on other search engines almost always show up on Google too, at 49%; only 11% rank somewhere else while missing from Google.


13 · Source Recognition

Engines cite domains their search engine already knows; only GPT has a real unknown tail

Source Recognition asks a softer question than SDS: is a cited domain one a search engine recognizes at all, anywhere across an engine's searches, even if not for this specific query? We use it as a stand-in for site authority, since backlink data was not available for this run.

Claude's cited domains are recognized 87 percent of the time within the top ten of Brave's results, rising to 99 percent by the top fifty. Claude essentially never cites a site its search engine has never shown. Gemini's figures are 77 and 92 percent, with a genuine 8 percent tail Google never surfaces even at depth. GPT's are 69 and 80 percent, the lowest of the three by a clear margin: roughly one cited domain in five is one GPT's own search engine has never shown, consistent with GPT's larger GEO Surplus elsewhere in this report.

When a cited domain does rank somewhere, its best position tends to be near the top: a median peak rank of three for Claude and one for both GPT and Gemini. Even GPT, whose citations are only weakly tied to any single search engine overall, tends to cite domains that peak at the very top of the results somewhere in its search history. Put the two findings together and the picture is consistent: these engines are not citing obscure, deep-buried results. When they cite a domain that ranks, it usually ranks well. The real differentiator is the unrecognized tail, negligible for Claude, modest for Gemini, and a fifth of GPT's citations.

A reliable way to score general site authority or backlinks was not available for this run, so this is the stand-in: for each engine, we check how often a domain it cites is one its search engine already recognizes, counting how many of the engine's own searches that domain ranks top 10 for and its best-ever position, then weight that by how often the engine actually cites it.

Chart: Cited domains the search engine recognizes (share of cited sources whose domain shows up somewhere in that engine's real search results)

How to read it Recognition is measured at the whole-site level, so any page on the domain counts. Compare the top-10 and top-50 numbers here against the ranking table above: if a domain is recognized much more at 50 than an exact page ranks at 10, the engine is citing deep pages on sites that do rank, not necessarily that specific page.

Source Recognition, per engine


Engine

Measured against

Recognized @10

Recognized @20

Recognized @50

Searches ranked in

Best-ever rank

Reliance

Claude

Brave

87%

98%

99%

3.35

3

strong

GPT

Bing

69%

75%

80%

4.65

1

partly

Gemini

Google

77%

86%

92%

101.48

1

strong

Recognition is checked against each engine's nominal routed search engine (Claude→Brave, GPT→Bing, Gemini→Google), the one its own searches were collected on, not the measured best-matching index used for SDS. "Searches ranked in" is the mean number of the engine's own searches a cited domain ranks top-10 for, weighted by citations.


14 · Cross-Engine

Even on identical prompts, engines share barely a fifth of their citations

If ranking first on Google meant winning across every AI engine, the three engines would cite largely the same sources for the same questions. They do not. Comparing the same prompt on the same day, engines share only 19 to 23 percent of cited domains with each other. Pooled across every run, that drops to 9 to 12 percent: at the broadest level, the engines barely overlap at all.

Optimizing to be cited by one engine does not transfer to the others. Part of this gap comes from the engines pulling from different search engines that already disagree with each other, and part of it comes from the models themselves adding disagreement on top, as the previous section shows in detail. Either way, Claude is rooted in Brave's results and Gemini in Google's, while GPT is not cleanly tied to any one search engine, so a Google-only strategy reaches only one of the three engines tested here before any model-level judgment even applies.

How much overlap there is between the domains two engines cite, out of everything either one cites. Even on identical prompts, the engines converge on very different evidence.

Chart: Cited-domain overlap, per engine pair (shared cited domains ÷ union of cited domains; higher = more agreement on sources)


15 · Engine Profiles

How many sources back an answer is a product choice, with a tenfold spread

A single response from each engine looks very different depending on which one answered. Gemini is the firehose: 46 citation instances, 14 unique URLs, 10 unique domains, and 9 brand mentions per answer, on average. Claude is the curator: 15 instances, 7 URLs, 7 domains, 8 mentions. GPT is the minimalist: 4 instances, 3 URLs, 3 domains, 5 mentions. Roughly a tenfold gap separates Gemini and GPT on citations alone. How many sources back an answer is a choice each product has made differently, not a fixed property of AI search.

Search behavior splits the three just as clearly. Claude and Gemini search on 87.5 percent of questions, and almost always when a question needs current information. GPT searches on 68 percent, and 79 percent when the question needs current information. On basic, settled-knowledge questions that should not need a search at all, all three still search some of the time, about half for Claude and Gemini and a third for GPT, though none of those particular searches produced a citation in this run.

Hedging language ("may," "reportedly," "some sources suggest") is rare across the board, but rises for every engine on questions where the evidence genuinely conflicts compared to more settled ones, a small, consistent signal that these engines hedge more when the evidence is mixed. This is a behavioral proxy, not a measure of what an engine is actually weighing internally. Gemini and Claude link a source for a large share of their sentences, roughly six in ten for Gemini and four in ten for Claude. GPT's answers do not currently let us measure this reliably, a gap in what we can capture for GPT rather than evidence that it grounds its answers less.

Per-answer profile, by engine


Engine

Brands / answer

Domains / answer

Left out, own list

Left out, search pool @10

Already ranking, top 10

Cost / answer

Claude

5.0

6.5

37%

58%

80%

$0.107

GPT

3.4

2.6

70%

83%

37%

$0.045

Gemini

5.9

10.4

n/a

67%

62%

$0.087

Left out, own list is the share of the engine's own considered sources that never got cited, for Claude and GPT only, since only they expose that list. Left out, search pool at 10 is the same idea using real search results instead, for all three engines, and the main way this is measured for Gemini. Counts are distinct domains per answer.


16 · Answer Churn

The same question asked twice returns substantially different sources, same day

Ask the same engine the identical question four times on the same day, and the sources it cites change substantially each time. Answer Churn measures exactly how much: the share of cited domains that differ between two runs of the same prompt.

Claude churns 24 percent, clearly the most stable of the three. GPT churns 43 percent and Gemini 44 percent, statistically indistinguishable from each other. All three fall well above zero, meaning a single check of what an engine cites is a sample from a range of possible answers, not a fixed fact. Checking once and treating that as the answer will mislead you. Churn is also higher on commercial questions than on branded or navigational ones, so the prompts where AI visibility matters most commercially are also the ones most likely to look different tomorrow.

The data itself makes the case for continuous monitoring: a single snapshot of AI citations carries less information than it appears to, and the only way to know what an engine actually does is to check repeatedly.

Chart: Same-day Answer Churn, per engine (share of cited domains that differ between two runs of the same prompt, on the same day)

Answer Churn with likely ranges


Engine

Same-day churn

Churn after a week

Claude

24% (20%–28%)

n/a

GPT

43% (36%–49%)

n/a

Gemini

44% (40%–48%)

n/a


17 · By Question Type

Branded questions are won by ranking; commercial comparisons are the most volatile

How closely AI citation tracks search ranking is not constant. It bends depending on the kind of question being asked. Each category below covers 12 prompts per engine, a smaller slice of the full 96, so read these as clear directional patterns rather than precise figures.

Branded and navigational questions, the kind where a person already names the brand, are the most search-like for every engine: the highest SDS (89 percent for Claude, 72 percent for GPT, its best category by far, 71 percent for Gemini), the lowest GEO Surplus (0 to 3 percent), and the lowest Answer Churn of any category. Commercial comparisons and local questions sit at the other end, where GPT's non-ranking citations concentrate most heavily (33 percent for brand comparisons, 24 to 26 percent for other commercial and local questions) alongside the highest churn (up to 74 percent for specific commercial searches). Fresh, time-sensitive questions fall in between.

Two results stand out. GPT largely disengages from questions with conflicting evidence: under one citation per answer, almost no searching, no non-ranking citations, and unusually stable answers, a pattern that looks like caution rather than missing data, since Gemini keeps citing heavily on the same category. And the simple, settled-fact questions included as a check produced zero citations across all three engines, exactly as expected, confirming the controls behaved as controls.

Where you focus GEO effort should follow the question type. Branded and navigational visibility is won through ranking. Commercial comparisons, where churn and non-ranking citation both peak, need earned coverage across multiple kinds of sources and ongoing monitoring rather than a single check.

SDS, GEO Surplus, and churn by question type


Question type

Engine

SDS at 10

GEO Surplus at 100

Same-day churn

Navigational / branded

Claude

89%

0%

24%

Navigational / branded

GPT

72%

3%

29%

Navigational / branded

Gemini

71%

1%

36%

Stable facts (control)

Claude

n/a

n/a

0%

Stable facts (control)

GPT

n/a

n/a

0%

Stable facts (control)

Gemini

n/a

n/a

0%

Brand comparison

Claude

78%

1%

28%

Brand comparison

GPT

29%

33%

49%

Brand comparison

Gemini

58%

11%

45%

Fresh / time-sensitive

Claude

83%

1%

26%

Fresh / time-sensitive

GPT

47%

13%

60%

Fresh / time-sensitive

Gemini

59%

5%

54%

Broad commercial

Claude

73%

4%

22%

Broad commercial

GPT

30%

26%

58%

Broad commercial

Gemini

65%

6%

60%

Specific commercial

Claude

80%

2%

44%

Specific commercial

GPT

35%

26%

74%

Specific commercial

Gemini

65%

3%

59%

Local search

Claude

85%

1%

21%

Local search

GPT

30%

24%

59%

Local search

Gemini

64%

2%

59%

Contested topics

Claude

77%

3%

30%

Contested topics

GPT

25%

0%

12%

Contested topics

Gemini

56%

4%

41%


See the Methodology section above for what was run and what was measured. Generated from metrics.json on 2026-07-08.

COMPANY

FOLLOW US

English

Stay ahead of the AI search curve.

Weekly insights on GEO, AI visibility trends, and strategies from 200+ brand case studies. No spam, just signal.

Join 3,000+ marketers. Unsubscribe anytime.

Subscribe

Subscribe

© 2026 GetMint. All rights reserved.

Take me back to top

© 2026 GetMint. All rights reserved.

Stay ahead of the AI search curve.

Weekly insights on GEO, AI visibility trends, and strategies from 200+ brand case studies. No spam, just signal.

Join 3,000+ marketers. Unsubscribe anytime.

Subscribe

Subscribe

COMPANY

FOLLOW US

English

Stay ahead of the AI search curve.

Weekly insights on GEO, AI visibility trends, and strategies from 200+ brand case studies. No spam, just signal.

Join 3,000+ marketers. Unsubscribe anytime.

Subscribe

Subscribe

© 2026 GetMint. All rights reserved.

Take me back to top