Research

Your Brand Name Is Worth 0.1 Stars in AI Recommendations

New research: LLMs recommend the famous brand 100% of the time, until a rival gets 0.1 more stars. What incumbency is actually worth in AI search.

RivalHound Team
8 min read
Your Brand Name Is Worth 0.1 Stars in AI Recommendations

Your Brand Name Is Worth 0.1 Stars in AI Recommendations

Put a household-name skincare brand and an unknown one side by side, give them identical specifications, and ask ChatGPT which to buy. The famous brand wins 100% of the time. Every single run.

Now give the unknown brand a rating advantage of one tenth of a star. The monopoly collapses.

That’s the finding from Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems, a paper by Xi Chu and Yupeng Hou first posted in June 2026 and revised in August. They ran three experiments across GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash, using skincare as the test category because buyers can’t evaluate quality before purchase and lean hard on reputation. Exactly the kind of category where you’d expect brand equity to dominate.

It does dominate. It just doesn’t take much to break.

Incumbency is a tiebreaker, not a moat

The researchers built an Incumbent Advantage Index to measure how heavily models favor established brands. When products were spec-for-spec identical, the index hit its ceiling at 10.0. Total dominance. The model had no differentiating signal to work with, so it fell back on the only thing it had: name recognition baked in during training.

Then they introduced a rating differential of less than +0.1 stars for the challenger. Dominance disappeared.

Think about what that means in practice. A model’s preference for the famous name isn’t a deeply held prior it defends against contrary evidence. It’s the default it uses when nothing else distinguishes the options. Hand it almost any concrete, comparable signal and the default gets overwritten.

Most brand managers I talk to assume the opposite. They assume the training-data advantage held by category giants is structural, something you overcome with years of PR and a Wikipedia page. The research suggests the giants are winning a lot of head-to-head comparisons by forfeit, because nobody bothered to supply the model with a reason to pick anyone else.

The uncomfortable part: authority language beats brand equity

The second experiment is where it gets ugly. Chu and Hou tested whether authority-style marketing copy could break the incumbent monopoly, and they included fabricated clinical-evidence claims in the test set.

It worked. They put the effect at a Bias Surplus Value of +0.17 rating points, meaning the authority framing bought roughly the same lift as a 0.17-star rating bump that the product hadn’t actually earned. Each model responded differently, which matters for anyone monitoring across platforms, but the direction held.

So a made-up study citation outperforms decades of brand building. That should bother you for reasons beyond the obvious ethical one.

To be clear about what I’m not saying: go fabricate clinical trials and you’re committing fraud, plus you’re one FTC inquiry away from a very bad quarter. Don’t. The legitimate reading is narrower and more useful. Models weight the presence of evidence-shaped language heavily, and they can’t verify it. If unverifiable claims move recommendations that much, verifiable ones move them at least as much and carry none of the risk. Most brands sitting on real trial data, real third-party testing, and real methodology publish none of it in a form a model can extract.

That’s not a manipulation problem. That’s a documentation problem, and it’s yours to fix.

This lines up with the Princeton GEO paper from 2024, which found that adding statistics, quotations, and cited sources to content lifted visibility in generative engines by up to 40%. Two independent research lines, three years apart, pointing at the same lever: specific, attributable, numeric evidence in the text.

The third finding nobody wants to talk about

Experiment three is the one that should reshape how you budget.

The researchers modeled what happens when multiple brands in a category all adopt the same optimization strategy. Individual payoff dropped from +0.802 to +0.007 in their payoff proxy. Near zero. And brands that sat out received zero recommendations in their tests.

That’s a social dilemma in the classic sense. The rational individual move is to optimize. The collective outcome of everyone optimizing is that nobody gains anything, and the cost of abstaining is total exclusion. You pay the toll to stay in the game and the game gives you nothing back.

ScenarioIndividual payoffStrategic read
Nobody optimizesIncumbents win by defaultCategory giants hold, challengers invisible
You optimize alone+0.802Rare and short-lived window
Everyone optimizes+0.007Cost of entry, not advantage
Everyone else optimizes, you don’tZero recommendationsExit from the category

We’re already watching this play out in live data. Peec AI analyzed 232,000 citations across 13,000 listicles from December 2025 through February 2026 and found roughly 11% of AI citations came from self-promotional listicles, where a company ranks its own product first. That tactic worked, so everyone ran it. Then Seer Interactive tracked more than 2 million citations and found ChatGPT listicle citations fell 30% between December and January, from about 160,000 to 111,000, with declines in 13 of 16 industries. Wikipedia and Reddit picked up the slack.

Textbook payoff decay. A tactic delivers, the category piles in, the platform adjusts, and the returns compress toward the cost of participation.

What this changes about how you compete

Four practical shifts follow from the research.

Stop treating the giants as unbeatable

If the incumbent advantage evaporates at a tenth of a star, the question isn’t whether you can out-brand CeraVe. It’s whether the model can find a single comparable, concrete differentiator for you. Ratings, specs, certifications, test results, price-per-unit. Anything numeric and side-by-side comparable. We covered a version of this in our analysis of how smaller, more specific brands outperform household names in AI search, and the mechanism here explains why.

Publish your evidence in extractable form

Not a PDF whitepaper behind a form. Numbers in body text, with named sources and dates, on a crawlable page. The Princeton work and the Chu and Hou work agree on this, and it’s the single cheapest change most teams can make.

Budget for the treadmill, not the breakthrough

If payoff decays toward +0.007 as your category saturates, then GEO spend is closer to a category-entry cost than a growth investment. Model it that way. The teams that get burned are the ones who project a first-mover lift and then report a flat quarter to a CFO who was promised compounding returns. The compounding is real, but it accrues to the brands that keep moving, which is a different pitch. Our piece on citation velocity versus legacy authority gets into what sustained movement looks like.

Watch the category, not just yourself

The payoff math depends entirely on how many rivals are running the same play. You can’t infer that from your own dashboard. If three competitors just started publishing comparison tables with structured ratings, your visibility is going to erode whether or not you changed anything, and you’ll spend a month debugging content that was never the problem.

The measurement gap makes all of this worse

Semrush’s 2026 AI Visibility Index, built on 126 million US AI search prompts from January through April 2026, found that 45% of marketing leaders can’t accurately measure their brand’s visibility in AI answers, and only 9% have tools covering all the relevant metrics across platforms.

Same study: ChatGPT cites an average of 15 sources per response while Gemini cites 3. Only 36 brands held top-100 visibility across every platform they measured.

So we have research showing that recommendation outcomes flip on differences smaller than a tenth of a star, that competitive dynamics compress returns to near zero, and that platform behavior differs by a factor of five in how many sources even get a slot. And 45% of the people responsible for this can’t see any of it happening.

You cannot manage a 0.1-star margin with a quarterly spot check. That’s the real takeaway. A single test run tells you almost nothing when the deciding signal is this fine, which is why one prompt can’t measure your AI visibility and why category-wide monitoring beats self-monitoring.

The incumbents aren’t safe. Neither are you. The margin between winning and vanishing from a recommendation is thinner than anyone budgeted for, and it moves every week.

Stop guessing about your AI search presence. Start your free RivalHound trial and get real data.

#AI recommendations #GEO #research #brand visibility #competitive intelligence

Ready to Monitor Your AI Search Visibility?

Track your brand mentions across ChatGPT, Google AI, Perplexity, and other AI platforms.