AI Cited Your Page. That Doesn't Mean It Used a Word of It.
Perplexity lists 16 sources per answer, ChatGPT 7. Then ChatGPT draws on its sources 4x more heavily. A 21,000-citation study on why citation counts lie.
AI Cited Your Page. That Doesn’t Mean It Used a Word of It.
Ask Perplexity a buying question and it hands back an answer with about 16 sources attached. Ask ChatGPT the same thing and you get about 7. Every AI visibility dashboard on the market, ours included, would score the Perplexity answer as the bigger opportunity. More slots, more chances to be in one.
Then measure how much each cited page actually shaped the words in the answer, and the ranking flips. ChatGPT’s cited pages score roughly four times higher on influence than Perplexity’s or Google’s. Perplexity’s list is long because most of what’s on it is decoration.
That finding comes from “From Citation Selection to Citation Absorption”, a paper posted to arXiv on April 28 by Zhang Kai, He Xinyue, and Yao Jingang, three independent researchers in China. It’s not peer reviewed, the sample is 602 prompts, and the whole thing is a single-day snapshot. I’ll get to the caveats. But it’s the first study I’ve seen that tries to separate two things every citation tracker mashes together: whether you got picked, and whether you got used.
Selection and absorption are different games
The paper splits the citation process into two stages. Selection is whether the engine triggers a search and puts your page in its source list. Absorption is how much your page then contributes to the generated text.
To measure absorption, the authors built an influence score from 0 to 1 for every cited page they could fetch. It’s a weighted blend of five signals: how many times the page is referenced in the answer (capped at three), whether it appears later in the answer rather than only at the top, how many paragraphs of the answer it touches, TF-IDF cosine similarity between page and answer, and bigram-plus-trigram overlap between the two. In plain terms, did the answer reuse your language, your numbers, and your structure, or did it just link to you.
Across ChatGPT, Google AI Overviews with Gemini, and Perplexity, the 602 prompts produced 21,143 valid citations. The authors fetched 18,151 of the cited pages and scored them.
| Platform | Mean citations per answer | Mean influence per cited page |
|---|---|---|
| ChatGPT | 6.88 | 0.2713 |
| Google AI Overview / Gemini | 12.06 | 0.0584 |
| Perplexity | 16.35 | 0.0646 |
Read the two columns against each other. Perplexity cites 2.4 times as many pages as ChatGPT and draws on each one about a quarter as much. Google sits in the same place as Perplexity on absorption despite citing fewer sources. ChatGPT is the outlier: a short list, heavily used.
This matches what you’d expect from how the products are built. Perplexity’s whole pitch is showing its work, so it attaches sources generously. ChatGPT, which runs a sliding window over a handful of pages and synthesizes, leans hard on the few it actually reads. Neither is wrong. But they mean a Perplexity citation and a ChatGPT citation are not the same unit, and adding them up in one visibility score is counting apples and pictures of apples as the same fruit.
What a page that gets used looks like
The more useful part of the paper is the contrast between pages in the top influence quartile and pages in the bottom one.
| Page feature | Top quartile vs. bottom quartile |
|---|---|
| Word count | 11.44x higher |
| Heading density | 12.50x higher |
| List density | 8.94x higher |
| Paragraph count | 5.69x higher |
| Semantic similarity to the answer | 2.31x higher |
Those are big multiples. The pages the models actually pull language from are long, chopped into many headed sections, and full of lists. They’re reference documents, not landing pages.
Content type mattered too. Pages containing code blocks scored 76.88% higher on influence than pages without. Pages with explicit definitions scored 57.33% higher. Pages with comparisons, 55.28% higher. The authors’ phrase for this is “evidence containerization”: a definition, a number with a date on it, a side-by-side table, a numbered procedure. Each one is a self-contained unit the model can lift without reconstructing context around it.
That lines up with a separate finding reported in the July 2026 survey of GEO research I covered last week. A forthcoming SIGIR paper ran 252,000 trials across six models and found explicit prices and recent dates moved citation, while formatting changes on their own had weak effects. Two studies, different methods, same shape. The model wants facts it can pick up whole. Bolding your headers doesn’t give it any.
The FAQ finding nobody will like
Every GEO checklist published since 2024 tells you to convert pages into question-and-answer format. The logic sounds airtight: models answer questions, so hand them pre-answered questions.
In this dataset, Q&A pages scored a mean influence of 0.0947 against 0.1005 for everything else. That’s 5.74% lower. Not a collapse, but the wrong direction, and the authors are blunt about it: “do not treat FAQ conversion as a universal GEO intervention.”
Their guess at the cause is that a lot of FAQ pages are thin. Ten questions, two sentences each, no numbers, no comparisons. A page like that is easy to select, because the question text matches the query, and useless to absorb, because there’s nothing in it worth quoting. It gets the citation and contributes nothing. Which is exactly the failure mode this paper exists to name.
I’d put it this way: the FAQ format isn’t the problem, the FAQ content usually is. A question header followed by 400 words with a definition, a figure, and a comparison is an evidence container that happens to start with a question mark. A question header followed by “Yes, Acme supports SSO. Contact sales to learn more” is a citation you’ll count and a reader you’ll never get.
Who gets picked, and who gets read
The source-type breakdown explains why so many brands feel invisible despite decent coverage.
| Source type | ChatGPT | Perplexity | |
|---|---|---|---|
| Official sites | 34.22% | 46.35% | 44.07% |
| News media | 31.17% | 18.99% | 16.07% |
| Vertical publishers | 22.13% | 22.00% | 18.99% |
| Combined | 87.52% | 87.34% | 79.12% |
Three source types take 79% to 88% of citations on every platform. That’s the selection stage, and it’s the concentration we’ve seen in every large citation study, including the 15 domains that carry most AI citations.
Now overlay absorption. Across the whole dataset, encyclopedia pages averaged 0.2144 influence, academic publishers 0.1118, commercial sites 0.1028. News media averaged 0.0726, near the bottom.
So news gets nearly a third of ChatGPT’s citation slots and then contributes less per page than almost any other type. That’s not a contradiction. A news article is a good retrieval match (fresh, relevant, authoritative domain) and a weak absorption source (narrative prose, few definitions, numbers scattered through quotes). The model picks it and then paraphrases around it. Meanwhile an encyclopedia entry or a dense commercial reference page gets its actual sentences reused.
For a brand, the practical split is this. Earned media gets you into the list. Your own reference content, if it’s built right, is what gets read once you’re there. Most teams invest in the first and let the second rot.
Prompt wording changes the citation set, again
A small side experiment in the paper is worth flagging for anyone running a monitoring program. The authors reran 60 prompts in three phrasings: natural, with an explicit request for sources, and with an expert-role instruction.
| Phrasing | ChatGPT citations | Google citations | Perplexity citations |
|---|---|---|---|
| Natural | 7.30 | 14.05 | 15.70 |
| Explicit source request | 6.15 | 15.90 | 17.15 |
| Expert role | 7.95 | 10.40 | 16.70 |
Telling Google to answer as an expert cut its citation count by a quarter. Asking ChatGPT to cite sources made it cite fewer. If your tracking tool’s prompts say “list your sources,” you’re measuring a different answer than your buyers see. We’ve argued before that one prompt can’t measure AI visibility. This is another reason: the citation count itself is prompt-sensitive, before you even get to which brands fill the slots.
Where the study is weak
I want to be specific about the caveats, because this paper will get quoted as gospel by people who never made it to the limitations section.
The influence score is a constructed proxy. The authors say so: it’s “a constructed observational proxy rather than a direct measure of hidden model attention.” It rewards lexical overlap, which means a page the model paraphrased heavily could score low while a page it quoted once scores higher. The 602 prompts were designed, not sampled from real traffic, and they were built to trigger search, which is why the search-trigger rates are 98.64% for ChatGPT, 99.67% for Google, and 100% for Perplexity. Real user prompts trigger search far less often. The same July survey cites a Schulte et al. study in which 57.8% of ChatGPT repetitions didn’t activate web search at all. About a quarter of cited pages (23.56%) couldn’t be fetched, and those failures probably aren’t random. Everything is descriptive. No causal claim survives.
What the paper does establish is narrower and still useful: citation count and citation influence are separable, they diverge sharply by platform, and the page features that predict influence are consistent and measurable.
What to do with it
Four moves, roughly in order.
- Check absorption on your own citations. Pull ten answers where you’re cited and read them next to your page. Does the answer reuse your number, your definition, your comparison? If it links to you and says nothing you said, that citation is doing less than your dashboard claims. This takes an hour and most teams have never done it.
- Audit your FAQ pages for evidence density. Keep the question headers. Replace two-sentence answers with a definition, a dated figure, and a comparison. If a question can’t support that, it probably shouldn’t be its own section.
- Build one long reference page per product category you care about. Many headings, lists, tables, defined terms, current prices where you can show them. This is the page type that scores in the top quartile. It’s also the page type most content audits flag as too long, which we’ve written about before.
- Split your platform expectations. On Perplexity, getting into the list is comparatively easy and worth comparatively little per slot. On ChatGPT, the list is short and each spot carries real weight. A visibility report that treats them as interchangeable is misreading both.
The competitive read
Absorption is also a sharper competitive signal than selection.
If a rival is cited alongside you on Perplexity, you’re both probably ornamental. If they’re cited on ChatGPT and the answer’s definition of the category is lifted from their glossary, they’ve won that query in a way a citation count never shows. The model is describing your market in their words.
That’s the thing to track. Not whether the competitor appears, but whether the answer sounds like them. When it does, go find the page it sounds like. It’ll be long, heavily headed, and full of definitions and numbers. Then build a better one.
RivalHound tracks your brand’s visibility across ChatGPT, Google AI, Perplexity, and more. Start monitoring to see where you stand.