Industry News

Cloudflare Split AI Traffic Into Three Lanes. The One That Buys Things Gets Blocked First.

On September 15, Cloudflare's new defaults block AI agents on ad-monetized pages while search bots stay welcome. The buying lane goes dark first.

RivalHound Team
8 min read
Cloudflare Split AI Traffic Into Three Lanes. The One That Buys Things Gets Blocked First.

Cloudflare split AI traffic into three lanes. The one that buys things gets blocked first.

OpenAI spent August teaching ChatGPT to shop. Product carousels landed in ChatGPT ads this month, letting several items from a retailer’s feed share one shoppable placement under a conversation (Digiday). The pitch to brands is simple: the agent is the new buyer, get in front of it.

Cloudflare spent the same month preparing to lock that buyer out. On September 15, new defaults take effect across its network: AI traffic gets sorted into three categories called Search, Agent, and Training, and for pages that display ads, two of the three get blocked by default (Cloudflare). Search stays welcome. Training gets the door. And Agent, the category that includes every bot acting on behalf of a person who might be about to spend money, gets the door too.

Most of the coverage has filed this under “publishers versus AI,” round twelve. That’s the boring reading. The interesting reading is what it does to the path between an AI recommendation and an actual purchase, because those now run through different lanes, and only one lane stays open.

The three lanes

Cloudflare’s taxonomy is worth quoting exactly, because the whole change hangs on these definitions. Search is “any behavior that collects or indexes your content, so it can answer questions about it later.” Agent is “automated behavior that is acting, usually in real time, on a person’s behalf, to get something done right now.” Training is “a crawler taking your content to train or fine-tune a model.”

Mapped to the bots you’d actually see in your logs:

LaneWhat it coversBots in this laneDefault on Sept 15 (ad pages)
SearchIndexing content to answer questions laterOAI-SearchBot, PerplexityBotAllowed
AgentReal-time action for a specific personChatGPT-User, Perplexity-User, browser agents like Comet and AtlasBlocked
TrainingHarvesting content to train modelsGPTBot, ClaudeBot, CCBotBlocked

OpenAI itself splits its traffic exactly this way, which is why the taxonomy works: GPTBot trains models, OAI-SearchBot builds the search index, and ChatGPT-User shows up when a person asks ChatGPT to go read a page right now. Under the new defaults, the first and third get blocked on ad-monetized pages while the middle one sails through.

The scope matters, so let me be precise about it. The September 15 defaults apply to new domains onboarding to Cloudflare, on pages that display ads. Existing customers keep their current settings, though Cloudflare says it will also update Training-crawler settings for existing customers unless they opt out before the deadline. So no, the web does not slam shut in two weeks. What changes is the default posture, and defaults are how the web actually gets configured. Cloudflare made training crawlers opt-in for new sites back in 2025, and most site owners never revisited the choice. Nobody un-flips a default they never knew existed.

The sentence nobody is reading closely enough

Buried in the announcement is a line that deserves more attention than the rest of the post combined:

“Multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).”

Read that again. Googlebot doesn’t separate its behaviors. The same crawler feeds Google Search, AI Overviews, and Gemini’s grounding, and Cloudflare has decided that crawlers refusing to declare separate lanes get judged by their worst one. If you block Training, you block Googlebot.

For the past year, blocking AI bots on Cloudflare was a one-click choice, and plenty of site owners took it and moved on. It felt free: keep the scrapers out, lose nothing. As of September 15, that legacy setting maps to a Training block, and a Training block now catches Google’s crawler itself. A site that set-and-forgot that toggle in 2025 could find itself invisible in the one search engine it never intended to leave. If you run on Cloudflare and you’ve ever touched the AI bot settings, checking the dashboard before September 15 isn’t optional housekeeping. It’s the difference between blocking GPTBot and deindexing yourself.

Why the agent lane is the one that matters

Blocking training crawlers is old news, and honestly, for most brands the visibility cost is modest and slow. Training data shapes what a model knows about you months later. You can argue either side of that trade.

The agent lane is different, because the agent isn’t researching. It’s transacting. When someone tells ChatGPT “find me a project management tool under $10 a seat and check whether it integrates with Slack,” ChatGPT-User goes and reads pages in real time on that person’s behalf. When someone uses Perplexity’s Comet browser to compare two vendors, the browser is the one doing the reading. These aren’t crawls that might influence some future answer. Each one is a live buyer, mid-decision, wearing a robot mask.

And this lane is growing faster than anything else on the web. Cloudflare’s own Q2 2026 data puts daily AI agent requests up more than 1,700% year over year (Converge Digest), part of the broader shift that pushed bots past humans in web traffic this June. Agents are the fastest-growing visitor to your site, and the September 15 defaults treat them as freeloaders to be turned away.

I understand why publishers cheer for this. The economics of AI crawling are genuinely lopsided: Cloudflare Radar data from June puts Anthropic’s crawl-to-refer ratio near 4,580 pages crawled per referral sent, OpenAI at 848 to one, against Google’s 5 to one (Nobori). If your business is selling ad impressions to human eyeballs, an agent that reads your review and relays the verdict without loading your ads is a parasite. Blocking it is rational.

But notice which pages the default protects: the ones that display ads. Those aren’t random pages. Ad-monetized pages are, overwhelmingly, third-party content. Review sites. Comparison roundups. Affiliate “best X for Y” listicles. Trade press. Forums. The entire evidence layer a buying agent would consult to validate a choice runs on ad revenue, which means the entire evidence layer is exactly what goes dark to agents first. Your own product pages don’t run display ads. They stay open.

Two forces, one direction

Put this next to what happened earlier in August and a pattern forms. On August 8, ChatGPT started running site: scoped searches at scale, and citations shifted hard away from Reddit, directories, and press toward brands’ own domains. That was the engine choosing to go direct to the source. Now the infrastructure layer is pushing the same way from the other side: third-party, ad-supported validation gets walled off from agents by default, while your own site remains the one place an agent can always go.

The strategic consequence lands the same either way. For the machine audience, your owned properties are becoming the primary, sometimes only, source of truth about your brand. The old GEO playbook leaned heavily on being praised in places you don’t control, and AI browsers were already ignoring most of what those pages look like. Now a growing share of agents won’t reach those pages at all. If your pricing, integrations, comparison arguments, and proof points don’t live on crawlable pages of your own domain, there is increasingly no second place for an agent to find them.

There’s a darker wrinkle for brands that depend on third-party validation. An agent that can’t reach the G2 listing or the trade-press review doesn’t tell its user “I was blocked, my answer is incomplete.” It answers with whatever it could reach. Sometimes that will be your site alone, which is fine for you. Sometimes it will be a competitor whose reviews live on a page without ads, or a stale cached snapshot from before the wall went up. The user just hears a confident answer. Nobody sees what was missing.

What to do before September 15

Five things, in order of urgency.

  1. If you’re on Cloudflare, open the AI traffic settings this week. Check whether the legacy “Block AI bots” toggle is on, and decide lane by lane. Blocking GPTBot is a defensible choice. Blocking ChatGPT-User means turning away live buyers, and inheriting a Googlebot block by accident is a self-inflicted wound.
  2. Separate your thinking about the three lanes. “Should we block AI?” was already a lazy question; as of September 15 it’s not even a coherent one. Search access drives whether you’re findable, agent access drives whether you’re buyable, and training access drives what models believe about you at a months-long lag. Our earlier breakdown of which bots to allow and which to block maps the specific user agents to each decision.
  3. Pull agent user-agents out of your server logs and baseline them now. ChatGPT-User, Perplexity-User, and the browser agents are your measure of how much live-buyer traffic is at stake before the defaults start reshaping it.
  4. Audit whether your own domain can carry the load. Pricing, integration lists, security posture, honest comparisons: on crawlable, unauthenticated pages, reachable without JavaScript acrobatics. The third-party pages that used to answer these questions for you are the ones going dark.
  5. Watch your citation sources for churn this fall. As agent blocking spreads through the default, expect AI answers to lean harder on whatever stays reachable. The brands that notice their evidence layer disappearing from AI answers in October will be the ones who were measuring in September.

The web is being renegotiated bot by bot, and the deals differ by lane. You don’t get a vote in Cloudflare’s defaults or OpenAI’s crawler architecture. You do get to know, day by day, whether the machines that recommend and buy on your customers’ behalf can still see you.

RivalHound tracks your brand’s visibility across ChatGPT, Google AI, Perplexity, and more. Start monitoring to see where you stand.

#Cloudflare #AI agents #AI crawlers #agentic commerce #GEO

Ready to Monitor Your AI Search Visibility?

Track your brand mentions across ChatGPT, Google AI, Perplexity, and other AI platforms.