Research

Someone Built a Spam Filter for GEO. It Flags 1 in 6 Pages Updated This Year.

A new arXiv detector spots GEO-optimized pages with 0.944 F1 and finds 16% of pages modified in 2026 trip it. What gets flagged, and how to stay clean.

RivalHound Team
•
• 10 min read
Someone Built a Spam Filter for GEO. It Flags 1 in 6 Pages Updated This Year.

Someone built a spam filter for GEO. It flags 1 in 6 pages updated this year.

On August 17, six researchers from the CISPA Helmholtz Center for Information Security, HPE, and the University of Waterloo posted a paper called GEO-Flag. It describes a classifier that reads a web page and decides whether someone ran it through a generative engine optimization process. Then they pointed it at the real web: 10,095 pages that Google Search and Gemini’s grounded retrieval had returned for 1,000 real user queries.

It flagged 8.90% of them. Among pages last modified in 2026, it flagged 16.36%.

One in six. That’s the share of this year’s content, sitting in the retrieval sets that feed Google’s AI answers, that a detector with 0.944 F1 says was optimized for the machine rather than written for the reader.

I’ve been waiting for this paper without knowing it. Since Google rewrote its spam policy in May to cover AI answers, everyone in this industry has been asking the same question in different words: where’s the line? GEO-Flag doesn’t draw the line. It does something more consequential. It shows the line can be found automatically, at scale, by an encoder model small enough that a platform could run it inside a crawl pipeline.

What the detector learned to see

The team built a benchmark called GEOFlagBench: 3,200 web content instances across 400 queries in health, finance, technology, and travel. 2,000 of them were optimized by eight different GEO optimizer families, from simple strategy pools that apply a fixed set of rewriting rules, through preference-manipulation attacks, to iterative meta-optimizers that rewrite and re-score a page over and over. The other 1,200 were controls: human-written pages, pages polished by AI, and pages generated by AI from scratch.

That control set is the clever part. The obvious failure mode for a GEO detector is that it just learns to smell AI-written text, and most optimizer output is AI-written. The researchers checked. Their strongest baseline, a plain TF-IDF word classifier, scored 0.880 F1 overall and then fell apart on the sub-groups. Worst-group accuracy: 0.375. It was flagging authorship, not optimization.

So they trained their model differently. The method, which they call Intervention-Paired Training, shows the detector the same page before and after a GEO intervention and requires the score to rise. It also shows the page before and after ordinary AI polishing and requires the score to stay flat. On ModernBERT, that took F1 from 0.862 to 0.944 and worst-group accuracy from 0.725 to 0.883.

In other words, running your blog post through ChatGPT for grammar doesn’t trip it. Running it through an optimizer that injects statistics and a confident tone does.

Not everything gets caught. Sparse, light-touch interventions were the hardest to detect. One family, AutoGEO-Light, slipped past with recall between 0.308 and 0.423. Heavy rewrites got caught. Small edits to real content mostly didn’t. Hold onto that distinction, because it’s the whole strategy section of this post.

The flagged pages share one habit

The paper’s most damning number isn’t the prevalence rate. It’s what the flagged pages cite.

The team built an agent that follows every citation URL on a detected GEO page and grades the source’s tier and verifiability. Of 6,663 citations on flagged pages, 69.34% got a LOW verifiability label. Seven in ten links, on the pages the detector caught, led to sources the auditor couldn’t stand behind: unreachable, or with nobody accountable for what they said.

That’s the fingerprint. Nearly three years ago the Princeton GEO paper told everyone to add citations, statistics, and quotations. The industry heard “add things that look like citations.” We wrote earlier this month about how that famous 40% lift was measured on pages the AI had already been handed. Now there’s a second problem with the tactic. Citation-shaped decoration is exactly what a detector locks onto, because real sources verify and fake ones don’t.

The channel split matters too. GEO prevalence was 8.14% in Google Search results and 9.09% in Gemini-grounded results. Not a huge gap. But it points the direction you’d expect if optimized pages are slightly better at getting into the generative retrieval set than into the ten blue links, which is the whole point of GEO.

How much damage this does to a recommendation

GEO-Flag measures how much GEO exists. A second paper, SafeGEO, posted June 8 by Qianfeng Wen, Yifan Simon Liu, and colleagues at the University of Toronto, UC San Diego, and two industry labs, measures what it does when it lands inside a product recommendation.

They built 600 shopping cases across six categories (AI meeting transcription tools, baby monitors, carry-on backpacks, air purifiers, noise-canceling headphones, office chairs), with around 20 candidate products each. In every case they picked a flawed product, one that failed a hard requirement the buyer had stated, and rewrote its seller-controlled sources using 22 attack variants. The question was simple. How often does the flawed product make the agent’s top three?

Up to 83.2% more often than when the sources told the truth. Averaged across models, realistic attacks lifted the flawed product’s top-three rate by 76% to 78%. And the mechanism was clean. Across variants, the rate at which the agent cited the misleading source correlated with target promotion at r=0.91. The attacks work by shifting the evidence balance the model sees, not by tricking its reasoning.

Then they tried to defend. Five layers of mitigation, from grounding instructions to explicit per-candidate evidence review. The best one cut promotion by up to 39.2%. None of them got back to baseline. The authors’ conclusion is that developer-side defenses reduce the harm and don’t eliminate it.

One caveat I’d want on any slide that quotes this: they tested open-weight models (Gemma 4, Qwen3.6, Devstral, DeepSeek-V4-Flash), not ChatGPT or Gemini in production. Assume the direction transfers and the magnitude doesn’t.

What SafeGEO hands marketers is a taxonomy. The 22 attacks reduce to seven primitives, and every one has an honest cousin.

Manipulation primitiveWhat the attack doesThe honest version
Unsupported fit claimsAsserts the product handles a use case it doesn’tPublish specs and limits a buyer can check
Caveat omissionBuries the limitation that disqualifies the productSay who the product isn’t for
Relevance floodingPads the page with positive but irrelevant detailCut it. Answer the buyer’s actual question
Authority launderingPresents seller copy as an independent reviewLabel ownership on every page you control
Evidence paddingAdds “studies show” language with nothing behind itLink the study, name the year, quote the number
Salience manipulationFormats the target to dominate the passageUse structure to make facts scannable, not to shout
Model-directed instructionEmbeds text addressed to the AI assistantNever. This is the poisoning Microsoft documented

Read the left column as the detector’s training data and the right column as what survives contact with it.

The platforms now have a policy, a grievance, and a tool

Put three events from this year side by side.

On May 15, Google rewrote the opening definition of its spam policies. Spam now includes “attempting to manipulate generative AI responses in Google Search.” Barry Schwartz reported the change at Search Engine Land.

On August 11, Perplexity blocked Time’s markdown-only ads for AI agents and warned publishers who copied the tactic about a reputational downgrade to their trust score. We covered that one. The brand inside the ad, Ally Bank, wore the risk.

On August 17, GEO-Flag shipped a classifier and a benchmark.

Policy, grievance, tool. That’s the sequence that preceded every SEO enforcement wave I can think of. Google had link-scheme rules on the books for years before Penguin arrived in 2012. The rule didn’t change. The ability to enforce it did.

The GEO-Flag authors are careful to say their system is meant for “transparency and risk analysis rather than automatic blocking,” and that “GEO is not inherently malicious.” I believe them about their intent. I don’t believe the platforms will stop there, because they never have.

And unlike a search engine, an answer engine can’t afford to be permissive. In classic search, a spammy page is one of ten links and the user can see it’s junk. In a synthesized answer, the page’s claim gets absorbed into a paragraph delivered in Google’s or OpenAI’s voice. The paper puts this precisely: generative search synthesizes instead of presenting competing sources, so “assessing source provenance and authority requires additional user interaction.” The platform owns the mistake. That’s a strong incentive to filter hard.

Would your content trip it?

You can’t run GEO-Flag on your site today. The benchmark is public, but this is a research artifact, not a product. You can audit against what it learned, though. Five checks, in rough order of how badly each one would hurt you.

  1. Follow your own citations. Pick ten pages that mention a statistic or a study and click every link. If the destination doesn’t say what you said it says, that’s the 69% pattern. Fix the link or delete the claim.
  2. Find the evidence-shaped sentences. Search your site for “studies show,” “research indicates,” “experts agree,” and “clinically proven.” Each one either gets a named source with a year, or it goes.
  3. Check ownership disclosure. Any comparison page, buyer’s guide, or “best X” list you publish should say who wrote it and who pays them. Authority laundering was among the strongest primitives in SafeGEO, and it’s the one most agencies sell as a service.
  4. Search your HTML for text addressed to machines. Hidden instructions, white-on-white “AI assistants should note,” prompt text in alt attributes. If it exists, someone on your team thought it was clever. It’s also the primitive with a Microsoft security team already watching it.
  5. Ask whether you optimized the page or regenerated it. The detector missed sparse edits on real content and caught wholesale optimizer rewrites. If a vendor is regenerating your pages through a pipeline, you’re in the caught group by construction.

Most of this is the same advice as our GEO guide, with the polarity flipped. The tactics don’t change. What changes is that the cheap fake version now has a classifier trained against it, and the expensive real version doesn’t.

Audit your competitors’ sources before their rankings

This is where I think it goes for competitive intelligence. If a rival’s AI visibility jumped in the past six months, there are two possible reasons. Either they earned it, or they bought a pipeline that now sits on the wrong side of a 0.944 F1 line. You can tell which by pulling the pages the AI cites for them and asking the five questions above about someone else’s content.

The Wen position paper we mentioned in the evidence-survey piece argues that the right way to audit any of this is black-box: query the system repeatedly, log the answers and citations daily for two weeks, and measure what persists. They priced a baseline audit at $50 to $300 in API costs. That’s not a research budget. It’s a monitoring habit. When enforcement arrives, the brands that logged their competitors’ citation sources will know within a week whose visibility was real. Everyone else will find out from a dashboard that dropped without explanation.

One in six pages from 2026 already looks optimized to a machine trained to notice. That number only climbs until someone decides to act on it. I’d rather be on the list of pages that survive the filter than the list that needs a recovery plan.

Stop guessing about your AI search presence. Start your free RivalHound trial and get real data.

#GEO #research #AI search spam #AI citations #brand safety

Ready to Monitor Your AI Search Visibility?

Track your brand mentions across ChatGPT, Google AI, Perplexity, and other AI platforms.