Research

Someone Read All 45 GEO Studies. The Famous 40% Lift Isn't What Vendors Say It Is.

A July 2026 arXiv survey of 45 GEO studies finds the Princeton 40% lift only applies to pages AI was already handed. No tactic has been shown to move traffic.

RivalHound Team
9 min read
Someone Read All 45 GEO Studies. The Famous 40% Lift Isn't What Vendors Say It Is.

Someone read all 45 GEO studies. The famous 40% lift isn’t what vendors say it is.

Every GEO pitch deck has the same slide. Add statistics, add quotations, cite your sources, and your visibility in AI answers goes up by as much as 40%. The number comes from a Princeton paper, it has been repeated for two and a half years, and it still headlines vendor decks and “GEO statistics” roundups.

On July 15, Olivier Martinez posted a survey to arXiv that reads every GEO study published since that Princeton paper. Forty-five of them, from November 16, 2023 through July 14, 2026. His verdict on the 40% figure is worth quoting in full: “This result does not mean that 40% more readers will click, nor that a page will gain 40% in retrieval probability. It means that, in this testbed, a source already provided to the generator receives a larger position-weighted share.”

Read that twice. The lift was measured on pages the AI had already been handed. Nobody checked whether the tactics help a page get found in the first place. And across all 45 studies, Martinez found no technique with “a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream clicks and conversions.”

I think this is the most useful thing written about GEO this year, and almost nobody in marketing will read it, because it’s dense with estimands and Jaccard scores. So this is the version for people who have to decide what to do on Monday.

Where the 40% actually comes from

The Princeton paper is Aggarwal et al., published at KDD 2024. The setup: take a query, put five documents in front of a generative engine, rewrite one of them using a tactic, and measure how much of the answer’s word count traces back to that document, weighted by how early it appears. The metric is called position-adjusted word count.

Quotation Addition moved that score from 19.3 to 27.2. That’s the 41% relative gain that got rounded to “up to 40%” and never rounded back down.

Three things about that experiment matter more than the headline number:

  • The document was one of five already in the context. Retrieval was fixed. The paper says nothing about whether a live engine would ever pull the page.
  • The metric is a share. If one of five sources gains, the other four lose. Martinez calls visibility in this setup “intrinsically relative and redistributive.”
  • No clicks, referrals, traffic, or purchases were observed. None.

So the honest sentence is this: if your page is already one of the handful an engine reads, adding a quotation makes it a bigger slice of the answer. That’s a real effect, and the survey grades it high-confidence. It’s just a much smaller claim than “GEO tactics boost visibility 40%.”

What the survey grades as proven, and what it doesn’t

Martinez breaks GEO into a seven-stage pipeline: whether search activates at all, whether your page is crawled and indexed, whether it’s retrieved into the candidate pool, whether it survives reranking and gets token budget, whether it’s cited, whether its facts get absorbed accurately, and whether a human clicks or buys. Most GEO evidence lives in stages four and five. Most GEO promises are about stages three and seven.

He then grades the main claims by how much evidence backs them. Compressed:

ClaimConfidenceThe caveat
A document already in context can change its rank or citationHighSays nothing about organic retrieval
Topical relevance and context position are the main determinantsHighMay shift with answer length and task
Engines differ from each other and change over timeHighHard to compare across products and periods
Extractable evidence (statistics, quotes, dates) helps a page get usedModerateDepends on intent, engine, and whether it’s true
White-hat GEO durably improves discoverabilityLowVery few end-to-end tests exist
Being cited predicts clicks or conversionsVery lowOne quasi-experiment, no causality established

The bottom two rows are the ones on the pitch deck. The top two are the ones with evidence.

The sharpest single data point against generic tactics comes from C-SEO Bench, a NeurIPS 2025 benchmark from Haritz Puerto and colleagues. They tested nine content-rewriting methods across six domains and 1,921 queries. Out of 54 method-and-domain combinations, three produced a statistically significant improvement. None in question answering. Many made ranking worse. And the thing that did work was old-fashioned: pushing the source higher in the model’s context, which is to say, ranking better at the retrieval step. They also reran the experiment at different adoption rates and found gains shrink as more competitors use the same tricks. Martinez’s phrase for that is “congested dynamics.” Mine is simpler. A tactic everyone can copy in an afternoon can’t be a moat.

The part vendors won’t quote

The survey spends a lot of pages on noise, and this is where it lines up with what we see in our own tracking.

Julius Schulte, Malte Bleeker, and Philipp Kaufmann’s Don’t Measure Once, from April, ran the same prompts across four engines for 45 days. Day-to-day overlap in cited sources came out at a Jaccard score of roughly 0.34 to 0.42. Repeat the query within 24 hours and you get similar churn. Other studies in the survey found 9% to 28% of decisions change even at temperature zero, and in one set of ChatGPT repetitions, 57.8% of runs never turned on web search at all. We hit the same wall when SparkToro and Gumshoe ran 2,961 prompts and got a repeated brand list less than 1% of the time.

Overlap across engines is worse. The survey cites 26% domain overlap between Bing Chat and Perplexity, and 53% of domains cited in Google AI Overviews sitting outside the organic top ten. We measured 11% shared domains between ChatGPT and Perplexity earlier this year.

Then there’s fidelity, which almost no vendor measures. Being cited doesn’t mean the answer says what your page says. Studies in the review put fully supported sentences at 51.5%, and citations that correctly back the sentence they’re attached to at 74.5%. About one atomic claim in nine was insufficiently supported. A separate April paper from Zhang Kai, He Xinyue, and Yao Jingang, From Citation Selection to Citation Absorption, ran 602 prompts through ChatGPT, Google, and Perplexity and found ChatGPT cites fewer pages but each one shapes the answer more. Your citation count and your actual influence on the answer are two different numbers.

Put those together and you get the survey’s blunt warning. The probability your page is cited equals the probability search fires, times the probability you’re retrieved given it fired, times the probability you’re cited given you’re retrieved. GEO tactics touch the last term. The first two are where most brands lose, and a great score on the last term hides that. In Martinez’s words, “high conditional citation rates can coexist with low commercial visibility.”

The first term alone is brutal. Profound’s analysis of about 730,000 ChatGPT conversations from late 2025 found roughly 18% of conversations trigger a web search, and the first turn cites sources 12.6% of the time versus 3.0% by turn 20. If search never fires, your quotation-optimized page is irrelevant. Nothing got retrieved.

The uncomfortable bit for us

RivalHound publishes a GEO guide that lists statistics, quotations, and citations as things to add. So does every competitor. The survey doesn’t say those are bad ideas. It says the evidence for them is conditional, the effect is relative, and nobody has connected them to traffic. That’s a different thing than “it works,” and I’d rather say so than keep quoting 40% with a straight face.

A second paper this summer makes a related point from the governance side. Yizhu Wen and colleagues, in a position paper accepted at ICML 2026, argue that offline GEO research and deployed engines are evaluated so differently that academia and industry have blind spots about each other. Lab benchmarks fix the retrieval step. Vendor dashboards fix nothing and report a single number. Neither tells you whether the tactic you paid for did anything.

What to actually do

Martinez’s own advice for practitioners fits in one sentence: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.” That last clause is the whole game. Six ways to act on it.

  1. Split your reporting into the three probabilities. Track whether your target prompts trigger search at all, whether your domain shows up in the cited set, and whether the answer text reflects what your page says. One blended “visibility score” hides which stage is failing. It’s the same problem we flagged with share of voice’s missing denominator.

  2. Stop measuring once. The survey’s minimum protocol is seven to eight repetitions per prompt, across several paraphrases, dates, and engines, with the spread reported. That starting count comes from Schulte’s data, and Martinez says your own pilot should set the real number. If a vendor shows you a point estimate with no variance, you don’t know whether the change was your content or Tuesday.

  3. Spend on retrieval before rewriting. C-SEO Bench found that ranking higher in the context beat every content trick, and the survey’s evidence table says the same thing: relevance and position dominate. Being in the candidate pool is the bottleneck. That means indexing, crawlability, and the third-party coverage that gets you into the grounding results to begin with.

  4. Keep the quotes and statistics, but only true ones. The moderate-confidence lever is extractable evidence that’s relevant and accurate. Borrowed or invented stats fail the survey’s own white-hat test (semantic preservation, evidentiary authenticity, no hidden instructions, disclosure). They also fail the Perplexity test, as Time and Ally Bank found out in August.

  5. Audit fidelity by hand. Pull a sample of answers that cite you and check whether the sentence attached to your citation is supported by your page. In one of the reviewed studies, a quarter of citations weren’t. A citation that misquotes you is a visibility problem wearing a win’s clothing.

  6. Ask vendors which stage their number measures. “Visibility up 30%” means nothing until you know whether that’s share of a fixed context, citation rate on live prompts, or referrals. Most tools report the middle one and let you assume the third.

The GEO research base is 45 papers and two and a half years old. Its most reproducible finding is that relevance and position win. Its most quoted finding is a share metric from a five-document sandbox. And its weakest link is the one that pays your salary. None of that means AI visibility isn’t worth working on. It means the work is retrieval and measurement, and the rewriting tricks are the cheap part everyone will copy.

Stop guessing about your AI search presence. Start your free RivalHound trial and get real data.

#GEO #research #AI citations #measurement #AI search

Ready to Monitor Your AI Search Visibility?

Track your brand mentions across ChatGPT, Google AI, Perplexity, and other AI platforms.