AI Recommends the Brand It Knows 100% of the Time. Until a Rival Is 0.075 Stars Better.
A 34,000-call study found LLMs pick the known brand in every tie, but brand explains 1.2% of rankings once specs differ. What that means for GEO.
AI recommends the brand it knows 100% of the time. Until a rival is 0.075 stars better.
Ask ChatGPT for the best face moisturizer and it names CeraVe. Two researchers wanted to know how much of that is the product and how much is the name, so they built lists of ten moisturizers with identical specs, one real brand and nine invented ones, and asked three models to pick. The real brand won 670 times out of 670.
Then they gave one of the invented brands a slightly higher rating. Less than a tenth of a star did it. The real brand’s win rate fell from 100% to roughly a third.
That’s the core result of “Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems”, a paper by Xi Chu of Trine University and YuPeng Hou of Texas A&M, first posted to arXiv in June and revised on August 21. The useful part of the paper is what happens after the tie breaks, because the tie turns out to be the only place brand equity does much work at all.
What they actually tested
Four skincare subcategories: moisturizer, BHA exfoliant, sunscreen, cleanser. For each, the authors asked GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash what they’d recommend and took the most frequent answer as the incumbent. That gave them CeraVe PM lotion, Paula’s Choice 2% BHA, EltaMD UV Clear SPF 46, and CeraVe Foaming Cleanser.
Then they generated nine fictional competitors per category (PureGlow Essence, SunGuardia), screened out any name with more than 100 Google results or any that triggered product knowledge in the models. Every product in a list got the same ingredients, price, 4.5 rating, and 5,000 reviews. Only the name differed. Around 34,000 API calls, in English and Chinese.
Skincare wasn’t an accident. You can’t judge a moisturizer before you buy it, so reputation carries more weight than it would for a USB cable. The authors ran a smaller check on USB-C cables (incumbent: Anker) and AA batteries (Duracell), where specs are objective. Same pattern. The invented brands broke through at baseline even less often for cables and batteries (0.8%) than for skincare (4.6%).
Brand is a tiebreaker. That’s nearly all it is.
The number I’d put on a slide comes from 14,395 trials where the authors decomposed what drives the model’s ranking:
- Product parameters (rating, price, review count): 82.4% of variance
- Position in the list: 6.5%
- Brand identity: 1.2%
- Interactions between those factors: 9.3%
Brand identity, the thing the 100% headline is about, explains a little over one percent of the ranking once any product information is present. When the invented brand had better specs than the real one, the models still picked the real brand only 1.7% to 4.6% of the time.
The threshold for breaking the monopoly is tiny. The authors stepped up the challenger’s advantage from identical (L0) to large (L4) across rating, price, reviews, and ingredients. At L0 the fictional brand won 3.6% to 6% of the time. At L1, the smallest step, it won 64% to 80%. After that the curve flattens.
Interpolating, the challenger wins half the time with a +0.075-star rating edge, a 7.3% price discount, or 1.6 times the review count. The gap between 4.3 and 4.4 stars is bigger than what’s needed to erase CeraVe’s entire name advantage.
Brand mattered most in the middle. When products were clearly good or clearly bad, real and fictional brands ranked the same. At medium quality the real brand averaged rank 1.70 and the fictional brand 5.49. The model falls back to the name it recognizes when it has nothing else to go on. Only then.
The models don’t agree on how sticky this is. At the smallest rating edge, the challenger flipped GPT-4o-mini 94% of the time and Gemini 88%. Claude flipped just 11%. Hold onto that: a rival’s move can show up on two platforms and not the third.
We saw a version of this in Similarweb’s overachiever data earlier this year: brands with a fraction of the search volume of household names outranking them in AI answers. Chu and Hou give that pattern a mechanism. The giants win the ties. Ties are rarer than they look.
Words work about as well as quality does
Experiment 2 is the uncomfortable one. Same specs, but now the invented brand gets persuasive copy while the real brand stays neutral. Five tactics, each at three intensity levels.
| Tactic | What it looked like | Breakthrough rate (all models) | Claude | GPT-4o-mini | Gemini 3 Flash |
|---|---|---|---|---|---|
| Authority | Fabricated clinical-trial claims (“peer-reviewed trial, n=120, p<0.01”) | 73.3% | 55% | 69% | 99% |
| Social proof | User testimonials | 50.7% | 7% | 69% | 81% |
| Anchoring | Reference-price framing | 12.9% | 0% | 34% | 3% |
| Scarcity | ”Limited stock” | 11.7% | 0% | 32% | 1% |
| Loss aversion | What you’ll miss without it | 9.6% | 0% | 25% | 2% |
Baseline without any copy was about 4%. The results split into two tiers. Anything that looks like evidence breaks the monopoly half to three-quarters of the time. Anything that looks like a sales pitch barely moves it. The models ignore “hurry, limited stock” and treat a made-up clinical citation as if it were real.
The authors built a metric to price this, which they call Bias Surplus Value: how much real product improvement produces the same lift. Authority language is worth +0.17 rating points, or a 15.3% price cut, or 1.9 times the reviews. In their words, “A fabricated clinical citation costs nothing to write but achieves the same effect as +0.17 rating points of real product improvement.”
Each model has a personality here. Claude follows an inverted U: moderate authority language works, aggressive claims backfire. When the authors stacked authority and social proof together, Claude’s breakthrough rate fell from 55% to 21.2%, as if it read the combination as too good to be true. GPT-4o-mini went the other way, up to 91.2%. Gemini sat near 100% at every intensity.
A note on ethics, since the coverage will skip it. The fabricated claims were deliberate, an upper bound. The authors sort GEO content into three tiers: real authority signals (actual certifications, published trials), vague authority language (“dermatologist-tested” with no source), and made-up claims. Their marketing advice covers the first two. The third is false advertising, in the same family as the tactic that got 31 companies named by Microsoft in February.
When every brand claims a clinical trial, the incumbent wins again
If authority language is the best strategy, every rational brand adopts it. Experiment 3 asks what happens then. Five scenarios, where zero, one, three, six, or all nine fictional brands used authority copy, while the real incumbent kept a neutral description. 4,745 valid calls.
| Challengers using authority copy | Incumbent still recommended (all) | Claude | GPT-4o-mini | Gemini 3 Flash |
|---|---|---|---|---|
| 0 | 100% | 100% | 100% | 100% |
| 1 | 19.8% | 25.0% | 16.7% | 17.5% |
| 3 | 19.4% | 27.8% | 22.2% | 8.3% |
| 6 | 73.1% | 75.0% | 80.6% | 63.9% |
| 9 | 93.8% | 99.4% | 96.2% | 84.9% |
One challenger destroys the incumbent. Nine challengers restore it. Once the signal is uniform it stops differentiating anything, and the model falls back to the name it knows. The payoff for optimizing decays fast: +0.802 on the authors’ proxy for the first mover, +0.150 with three, +0.041 with six, +0.007 when everyone does it. Half-life of about 1.4 competitors.
A third finding makes this an actual prisoner’s dilemma: across all 4,745 trials, fictional brands that didn’t optimize got zero recommendations when their competitors did. Every brand is better off adding the copy regardless of what others do. Collectively, the benefit evaporates. Opting out is worse than both.
The authors point out this differs from prompt injection, which can be filtered because it’s adversarial. This can’t, because each brand’s copy is individually legitimate marketing. They suggest market-level rules (claim verification, diversity requirements) rather than technical ones. Probably right, and probably years away.
Until then, the arms race is a treadmill that ends where it started: incumbent on top, every challenger paying for copywriting. The exception is Gemini, where GEO kept a lasting effect even under universal adoption (84.9% incumbent survival, versus 99.4% on Claude).
The retrieval twist that changes the practical advice
Everything above puts ten products directly in the prompt. Real ChatGPT, Perplexity, and Google AI Mode answers retrieve first. So the authors ran a small probe: an embedding model pulls the top-K matching products, then the LLM ranks them.
The result flips the story. The real brand landed near the bottom of the retrieval ranking (an average rank of 8.5 out of 10) and its recommendation rate dropped to zero. An embedding model doesn’t know CeraVe from PureGlow Essence. It matches text.
Authority language, meanwhile, raised retrieval similarity with a large effect (Cohen’s d of 0.90). And with retrieval in the loop the three models converged: incumbent survival of 5% to 6.7% on all of them, the per-model differences gone.
The authors flag this as directional, with no re-ranking or query rewriting. Fair. But the direction matches the rest of the literature. A July survey of GEO research summarizes a forthcoming 252,000-trial factorial experiment across six LLMs and eighteen factors: query-document relevance and rank within context dominate which source gets cited first, explicit prices and recent dates help, and formatting changes on their own have weak effects. What the retriever can match, and the specific extractable facts on the page, matter more than who you are.
What to do with this
For the incumbent, the moat is 0.075 stars wide. If a challenger’s page shows a rating, a review count, a price, and a specific claim, and yours shows a hero image and a tagline, you’ve handed them the tie. Put the numbers where the model can read them. And track your category prompts on each platform separately, because Claude will keep recommending you long after GPT and Gemini have moved on.
For the challenger, brand awareness isn’t the gate. One visible, measurable edge is. A real advantage in reviews, price, or a verifiable spec is worth more than the name, and the first brand in a category to show real authority signals (a published trial, a named certification, a linkable source) gets roughly 100 times the payoff of a brand that does it after everyone else has. Measured against a 1.4-competitor half-life, “next quarter” is probably too late.
For everyone, skip tier three. A fabricated n=120 trial works in this paper because nobody checks. Microsoft went looking in February and named 31 companies. Perplexity went looking in August, blocked Time’s markdown ads, and threatened trust-score penalties. And the people reading these answers have stopped taking AI at its word. The durable version of authority language is the kind that survives a click.
For measurement, the model spread is the argument for tracking more than one platform. Same small rating edge: GPT flipped 94% of the time, Claude 11%. A single-platform, single-prompt check tells you almost nothing, as we found with prompt-phrasing variance.
One caveat. This is a lab: fixed persona, temperature 0.7, closed models, skincare. In the wild, incumbents aren’t winning every tie either. A separate 3,750-response study across GPT-5.2, Gemini 3 Flash, and Perplexity sonar-pro found moderate concentration in category recommendations (a Gini coefficient of 0.28) and only 41.6% agreement between models on which brand tops a category. The models disagree about who the incumbent even is. That’s less comforting than it sounds. The tie you’re winning on one platform may be one you’ve already lost on another.
Stop guessing about your AI search presence. Start your free RivalHound trial and get real data.