Summarize this blog post with:
We analyzed 22,295 AI answers across ChatGPT, Perplexity, and Google AI Mode to measure how each engine distributes its citations across the five broad source families (brand websites and product pages, lists and comparisons and reviews, editorial and educational content, community and social platforms, and directories and marketplaces). The finding is that the three engines have fundamentally different source-family signatures. ChatGPT concentrates 69% of its citations in brand websites and product pages. Perplexity and Google AI Mode distribute their citations across three source families almost evenly, with brand websites at roughly 35%, lists at 33-34%, and editorial at 24-25%.
The dataset covers 115,843 citation events, 460 distinct B2B prompts, and 37 tracked organizations. Each citation was classified into a source family based on the page type of the cited URL. We report the share of each engine’s total citations that falls into each family. The samples per engine per family are in the thousands, so the family-level percentages are precise even when specific top-domain numbers are not.
Here are the questions we set out to answer:
- How does each of the three engines distribute its citations across source families?
- Which family does each engine over-index on relative to the other two?
- What does the source-family signature tell us about how each engine synthesizes answers?
- How does this stack with our earlier finding on which source families drive brand mentions?
- What does this mean for how a marketer should think about per-engine content strategy?
ChatGPT is a website-first citation engine. Nearly seven in ten of its citations come from brand websites and product pages, with the remaining three-tenths distributed thinly across lists (14%), editorial (12%), and community (5%). Perplexity and Google AI Mode are three-family engines. Both distribute their citations approximately evenly across websites, lists, and editorial, with a small share going to community sources and almost none going to directories. That structural difference means the same brand can face a very different citation landscape on each engine, and per-engine content strategy has to account for the specific family mix each engine actually reaches for.
Table of Contents
TL;DR
- What share of ChatGPT’s citations come from brand websites? 68.8%. Combined with lists (13.9%) and editorial (11.9%), those three families account for 94.6% of ChatGPT’s citations. Community sources sit at 5.3%. Directories are essentially absent at 0.13%.
- What share of Perplexity’s citations come from brand websites? 35.4%. Lists sit at 34.4%. Editorial at 25.2%. Community at 4.7%. Directories at 0.26%. The top three families are within ten percentage points of each other.
- What share of Google AI Mode’s citations come from brand websites? 37.9%. Lists at 33.1%. Editorial at 23.8%. Community at 4.8%. Directories at 0.46%. The distribution is close to Perplexity’s.
- What is the sample? Each engine’s citation total sits between 23,414 (ChatGPT) and 35,947 (Perplexity). The source-family shares are precise to within a percentage point.
- What should marketers do about this? Prioritize brand-website hygiene for ChatGPT visibility. Add investment in third-party lists and category editorial for Perplexity and Google AI Mode. Do not treat “AI citation strategy” as a single content playbook, since each engine reaches for a substantially different mix.
The Three Engines Have Fundamentally Different Source-Family Signatures
The engine-by-family share table.
| Source family | ChatGPT | Google AI Mode | Perplexity |
|---|---|---|---|
| Brand websites and product pages | 68.8% | 37.9% | 35.4% |
| Lists and comparisons and reviews | 13.9% | 33.1% | 34.4% |
| Editorial and educational | 11.9% | 23.8% | 25.2% |
| Community and social | 5.3% | 4.8% | 4.7% |
| Directories and marketplaces | 0.13% | 0.46% | 0.26% |
| Total | 100% | 100% | 100% |
ChatGPT’s citations are concentrated in one family. Nearly seven in ten come from brand websites or product pages. The remaining families each contribute low double-digit or single-digit percentages, and directories are effectively absent.
Perplexity and Google AI Mode look nothing like ChatGPT. Their top three source families (websites, lists, editorial) each contribute between 24% and 38% of total citations, and no single family dominates. The Perplexity distribution is 35% / 34% / 25% across the top three. The Google AI Mode distribution is 38% / 33% / 24%. Those two engines are much more balanced retrievers than ChatGPT is.
The community family is remarkably consistent across engines, at 4.7 to 5.3%. Community and social sources are a small but similar share of every engine’s citation mix. The directories family is remarkably small across all three engines. In our B2B panel, G2, Capterra, and similar directory pages contribute less than half a percent of total citations on any engine, even though we saw in the source-family-transfer piece that when they do get cited, they lift mention rate substantially.
Why ChatGPT Concentrates on Brand Websites
The mechanism behind ChatGPT’s website-heavy signature is that ChatGPT was built as a conversational model with parametric knowledge from its training data as the primary source of information. Live retrieval was added later as a supplementary feature. When ChatGPT does reach for a live web citation, it tends to reach for a canonical source of information about the specific entity being discussed, which in a B2B context means the brand’s own website or product page. The retrieval is confirmatory rather than exploratory.
That mechanism also explains why ChatGPT’s average answer is longer (3,361 characters) despite having fewer citations (3.6 per answer) than the other two engines, as we covered in the length-vs-citations piece. Most of a ChatGPT answer is prose synthesized from training data, with a handful of brand website citations providing verification for the specific entity mentions. The engine writes more, cites less, and when it cites, it cites the brand itself.
For a marketer, the practical implication is that ChatGPT visibility depends heavily on whether your brand’s own website is retrievable, indexable, and topically strong on the queries your buyers ask. If ChatGPT reaches for a brand website 69% of the time when it cites anything, and your website is either not indexed or not topically aligned with the buyer’s query, ChatGPT will either cite a competitor’s brand website or synthesize the answer entirely from training data without citing anyone. Both outcomes leave your brand out of the answer.
Why Perplexity and Google AI Mode Distribute More Evenly
Perplexity was designed as a citation-first synthesizer. Every claim in a Perplexity answer tries to have an inline source attribution, and the engine performs a fresh web search on every query. That design pushes Perplexity to retrieve a wider mix of source types, because the answer format requires source anchors for both branded claims (which pull brand-website citations) and category-level claims (which pull editorial and list citations).
Google AI Mode inherits from Google search infrastructure. Its retrieval draws on Google’s ranked results, which include the full mix of brand websites, editorial explainers, and list-style content. The 33% list share on Google AI Mode is particularly telling. Google’s organic ranking has long favored comparison articles and category rankings for commercial-intent queries, and Google AI Mode’s citation mix reflects that inheritance.
Both Perplexity and Google AI Mode look more like traditional search engines than like ChatGPT does, in the sense that both reach for a mix of source types that mirrors how the underlying web indexes those types. ChatGPT looks more like a conversational assistant that dips into the web for entity verification but relies on training data for the bulk of the answer.
For a marketer, the practical implication is that Perplexity and Google AI Mode visibility depends on a broader source portfolio than ChatGPT visibility does. You need brand-website hygiene like you do for ChatGPT, and you also need presence in third-party lists (G2 rankings, category comparison articles, top-10 blog posts) and editorial content (industry publications, explainer articles that cover your category). Optimizing only your own site will underperform on these two engines relative to ChatGPT because the source mix leaves so much room for third-party sources to displace brand-owned content.
Sources view showing source-family breakdown of citations for a tracked prompt panel across engines
How This Stacks With the Source-Family Transfer Finding
The current source-mix data extends the finding from the source-family transfer piece in a specific way. The transfer piece established that brand websites and directories are the two families where citations correlate with brand mentions (+15.8 pp and +23.9 pp lift respectively), while community, editorial, and list citations correlate with lower mention rates. The current data adds the retrieval-mix layer to that.
ChatGPT reaches for the mention-lifting family (brand websites) 69% of the time. Perplexity reaches for it 35% of the time. On the retrieval-mix logic alone, we would expect ChatGPT to produce higher brand mention rates than Perplexity, because the source-family mix ChatGPT retrieves from is the one that carries mentions.
The overall mention rates from our matched-comparison piece show the opposite. Perplexity mentions the focal brand 41.7% of the time on matched prompts. ChatGPT mentions it 32.3%. Perplexity has the higher mention rate despite reaching for the mention-lifting family less often.
The reconciliation is that Perplexity’s list-heavy citations (34% of the mix) surface many brands per answer. When Perplexity cites a “top 10 CRMs” article, the article names ten brands and any of those brands has a chance to appear in the answer text. When ChatGPT cites a specific brand’s own website (69% of the mix), that citation lifts the specific-brand mention rate but does not surface ten alternatives. The two mechanisms produce a similar aggregate mention rate through different retrieval and synthesis structures.
For a marketer, this means the relationship between citation-mix and mention rate is not additive. Getting more brand-website citations on ChatGPT lifts your specific-brand mention rate. Getting more list-inclusion mentions on Perplexity lifts your aggregate mention count because the lists themselves name multiple brands. The two are separate levers, and pulling only one leaves the other engine underoptimized.
How This Compares to Other Public Studies
Per-engine source preferences have been studied extensively in the last twelve months. The public studies converge on the direction of our finding while measuring at the domain level rather than the family level.
Profound analyzed 680 million citations across ChatGPT, Google AI Overviews, and Perplexity and reported that Wikipedia accounts for 47.9% of ChatGPT’s top-10 source share, versus much smaller shares on Google AI Overviews and Perplexity. Reddit is the dominant single source for Perplexity’s top 10 (46.7% share per a separate cut) and for Google AI Overviews. Their finding is at the top-domain concentration level. Our finding is at the source-family level. Both point at fundamentally different retrieval preferences per engine.
BrightEdge reported that “ChatGPT has almost no UGC presence (0.5%) and pulls heavily from government and .org domains.” Their UGC share is smaller than our 5.3% community-and-social share on ChatGPT, which likely reflects different classification methodology (UGC might exclude LinkedIn or Medium while our community-and-social family includes them). BrightEdge’s broader point about ChatGPT being “editorial and institutional” in tone is consistent with our finding that ChatGPT concentrates on formal brand-owned pages.
Leads Now AI reported that ChatGPT’s non-Wikipedia citations concentrate in established publications and review platforms like G2, and Perplexity’s citations concentrate on community discussion sources like Reddit. Their framing is that ChatGPT trusts “consensus reference material and established editorial brands” while Perplexity trusts “what practitioners and buyers say about you in public.” Our data is consistent with the first half of that framing but partly divergent on the second half. Perplexity in our B2B sample has a similar community share (4.7%) to ChatGPT (5.3%), so the “Perplexity is Reddit-heavy” framing is more accurate for consumer queries than for B2B queries.
Conductor analyzed 7 months of citation data and concluded that “every major AI engine has a persistent editorial identity” that differs by engine. Their finding at the philosophical level matches our finding at the numerical level. The engines are structurally different, and one content playbook cannot cover all three.
The pattern across these studies is consistent. Each AI engine has a distinct retrieval signature. Public research measures this at the top-domain level (Wikipedia versus Reddit versus G2). Our data adds the source-family level, which is more actionable for content strategy because a marketer can invest in a family of content types more easily than in a specific top domain.
What This Means for Per-Engine Content Strategy
Three things follow from the source-family mix data.
The first is that ChatGPT visibility depends on brand-website hygiene more than on any other single content investment. If ChatGPT reaches for brand websites in 69% of its citations, and your brand website fails on any of three fronts (unindexed by GPTBot, weak topical alignment on your category queries, or poor structure for AI extraction), ChatGPT’s citations will go elsewhere and your mention rate will suffer. The Content Optimizer can help audit whether your existing brand pages are the type of content ChatGPT actually retrieves for your tracked prompts.
The second is that Perplexity and Google AI Mode require a broader source portfolio than ChatGPT does. Brand-website content is roughly one third of what these engines retrieve. The other two thirds are lists and editorial. That means your visibility on Perplexity and Google AI Mode depends substantially on being present in third-party category rankings, comparison articles, and editorial explainers. Owned content alone will not carry your visibility on these two engines the way it can on ChatGPT.
The third is that community content is a small but consistent share of every engine’s citation mix (4.7 to 5.3%). Investment in Reddit, LinkedIn, and forum content will not move any engine’s citation share meaningfully, and by extension will not directly move brand mention rates. Community content builds category authority through the aggregate-signal mechanism Ahrefs measured (0.664 correlation between third-party brand mentions and AI citations), but it is not the retrieval mix any single engine is drawing from directly. The Citation Analytics view surfaces the specific source families per engine so you can see where your citations are actually coming from and adjust the content portfolio accordingly.
The Bigger Story
The industry has developed a simple mental model of AI citation strategy: get your content into the AI’s retrievable pool and it will get cited across all three engines. Our data contradicts that model. The three engines retrieve from fundamentally different source-family mixes, so a page that performs well as a citation source on ChatGPT may not be the type of page Perplexity or Google AI Mode reaches for. And vice versa. The 11% domain overlap between engines that Averi and Passionfruit reported reflects this at the top-domain level. Our data shows the same phenomenon operating at the source-family level.
The reframe is that AI citation strategy has to operate as a per-engine portfolio rather than a single content playbook. The commercial-content program (owned website, product pages, homepage) drives ChatGPT-specific citations. The list-inclusion program (category rankings, comparison articles, third-party review platforms) drives Perplexity and Google AI Mode citations. The editorial-inclusion program (industry publications, explainer articles) also drives Perplexity and Google AI Mode citations. Splitting the portfolio into three tracks and reporting each separately produces a strategy that matches how the engines actually behave.
This finding stacks with the other pieces in the State of AI Search series. Cross-engine consensus is rare, and source-family mix is one of the mechanisms that produces that low consensus. Brand websites and directories are the two families that drive brand mentions, and ChatGPT’s website-heavy signature translates that mention-driving mechanic into an engine-level advantage on specific-brand focus. Perplexity’s list-heavy signature translates the same underlying data into a multi-brand shortlist advantage that produces higher aggregate mention rates through a different mechanism. The engines are different products for source-portfolio purposes, and any strategy that treats them as one channel is optimizing against averages that hide the actionable variance.
The finding to hold onto from this piece is the simple one. ChatGPT’s citations are 69% brand websites. Perplexity and Google AI Mode split more evenly across websites (35-38%), lists (33-34%), and editorial (24-25%). Community content is a small share of every engine at 5%. Directories are less than half a percent everywhere but produce disproportionate mention-rate lifts when they do get cited. Build your content portfolio as three separate programs (commercial, list-inclusion, editorial-inclusion) and match each program to the engine that actually retrieves from it.
This research was conducted using Analyze AI, which tracks brand visibility, per-engine citation source-family mix, and page-type breakdowns across ChatGPT, Perplexity, Google AI Mode, and every other major AI engine.
Ernest
Ibrahim

![Featured image for ChatGPT Pulls 69% of Citations from Brand Sites [2026]](/_next/image?url=https%3A%2F%2Fwww.datocms-assets.com%2F164164%2F1785857882-15-source-family-mix.png&w=3840&q=75)

![Featured image for The State of AI Search in B2B [2026]](/_next/image?url=https%3A%2F%2Fwww.datocms-assets.com%2F164164%2F1785857899-16-state-of-ai-search-b2b-2026.png&w=3840&q=75)
![Featured image for Source Type Shifts AI Mention Rate 33 Points [2026]](/_next/image?url=https%3A%2F%2Fwww.datocms-assets.com%2F164164%2F1785857866-14-source-family-transfer.png&w=3840&q=75)
![Featured image for Top 10 Sites Cover Only 12% of AI Citations [2026 Study]](/_next/image?url=https%3A%2F%2Fwww.datocms-assets.com%2F164164%2F1785857692-05-top-domains.png&w=3840&q=75)