First-party research · September 2026

AI engines do not share a citation layer

Most advice about AI visibility treats ChatGPT, Perplexity, Claude and Gemini as one channel. Measured directly, they draw on substantially different sources — and when the same question is asked in another language, ChatGPT returns almost entirely different ones again.

What was measured

Matched buying-intent queries were run against all four engines with web retrieval forced on, and the hostname of every cited source was recorded. Two independent B2B verticals were tested: embedded insurance APIs, an emerging category with a thin content ecosystem, and payments infrastructure, a mature one with years of developer content and comparison sites behind it. Roughly 700 answers were recorded in total.

Finding 1: ChatGPT sits apart from the other three

For each pair of engines, the share of cited domains they had in common, expressed as a percentage of the smaller set:

Engine pairEmbedded insurancePayments
ChatGPT / Claude12%13%
ChatGPT / Gemini17%7%
ChatGPT / Perplexity29%20%
Perplexity / Claude44%32%
Perplexity / Gemini42%31%
Claude / Gemini48%30%
Distinct cited hostnames per engine, single scan per vertical, 11 queries × 3 runs.

The pattern holds in both markets. ChatGPT shares 7–29% of its sources with the others; the remaining three share 30–48% with each other. That it replicates in a thin emerging category and a mature content-dense one suggests it is a property of the engines rather than an artifact of how well a niche is covered.

Finding 2: language separates ChatGPT’s sources almost completely

The same query was asked in English and in a target language, with the request geo-targeted to a market where that language is spoken. Each figure is the mean of six comparisons — two queries, three runs each.

LanguageChatGPTPerplexityClaudeGemini
Turkish0%15%40%45%
Arabic4%9%36%40%
Spanish6%16%28%25%
Japanese12%4%29%35%
German32%48%52%71%
Share of English-query source domains also cited for the same query in the target language. 240 calls.

Turkish returned zero shared domains on every one of six samples. German is the clear exception, which fits a market whose business content sits close to the English-language web.

Per-cell ranges are wide — Claude’s Arabic figure averages 36% across a 6–73% spread — so these are bands rather than point estimates.

Finding 3: the consequence for a single company

Cover Genius, an embedded insurance provider, was measured across 11 buying-intent queries in its category. Its citation rate varied roughly fourfold depending on which engine was asked.

EngineScoreCitedMentioned
ChatGPT1512%21%
Perplexity2918%55%
Gemini5446%73%
Claude5548%71%
Same company, same queries, same week.

A second company in a different vertical showed the opposite failure. Checkout.com was named in answers by all four engines but had its own domain cited by only one — recognition without retrieval. Its score was 14 against Cover Genius’s 48, despite both being well-funded businesses in competitive categories.

Finding 4: the citation layer differs by market structure

In embedded insurance, the most-cited domains were the vendors themselves. In payments, they were comparison and review sites — comparepsp.com, paymentproviders.io, fintechspecs.com, thecfoclub.com — while vendor domains appeared less often than their brands did in the prose.

That difference matters more than the scores. In one market the work is getting your own content cited; in the other it is appearing on the intermediaries the engines trust. Advice that ignores market structure will point in the wrong direction half the time.

What this does not show

Two verticals and five languages, both chosen rather than sampled. Single scans per vertical for the engine-overlap figures. Scores for the same company have been observed to move by up to 16 points between measurements days apart, so individual numbers carry real uncertainty even where the patterns are consistent. Nothing here establishes a link between citation and revenue.

Method

ChatGPT was queried through the Responses API with web search forced on; without forcing it the model answers from memory and cites nothing, which reads as absence rather than as a failed search. Perplexity, Claude and Gemini used their native retrieval. Model versions were pinned. Answers returning zero sources were treated as failed calls and retried, never recorded as a zero. Every recorded answer is retained with its query, engine, run index, cited domains and full text.

Measured for your company

The same measurement runs weekly against any company in a covered category — your score per engine, the competitors named instead of you, and the consensus sources that decide the category. Pricing and access.

Full scoring method, weights and limitations: methodology. Questions or replication requests: [email protected]