Trending
August 10, 2026

Cloudflare's September 15 AI Crawler Defaults: Why a Training Block Can Take Googlebot With It

Cloudflare stops treating “AI bots” as one group next month. The headline is about making AI companies pay for content. The part that will quietly cost SEOs traffic is a single sentence about multi-purpose crawlers.

13 min read
Updated August 10, 2026

Key Takeaways

  • From September 15, 2026, Cloudflare splits crawler traffic into three categories — Search, Agent, and Training — each with its own block control.
  • The new default allows Search crawlers and blocks Agent and Training crawlers on pages that display ads.
  • Crawlers that do more than one job are governed by the most restrictive rule that applies — so blocking Training also blocks Googlebot, Bingbot, and Applebot on those pages.
  • The defaults apply to new domains, new sites added to existing accounts, and every site on the free tier.
  • You cannot signal your way out: Google has said it does not use llms.txt, and Cloudflare's Content Signals directive has no effect on any crawler. The dashboard setting is the only lever that acts.

What Happened: Cloudflare Split AI Crawler Traffic Into Three Categories

On July 1, 2026, Cloudflare announced that it would stop treating “AI bots” as a single category. In a post published for what the company calls Content Independence Day, Cloudflare split automated crawling into three distinct behaviours — Search, Agent, and Training — and gave every customer, including those on the free plan, an independent control for each one. The new defaults take effect on September 15, 2026.

The definitions matter, because they decide which switch your traffic lands under. Cloudflare defines Search as “any behavior that collects or indexes your content, so it can answer questions about it later.” Agent covers “automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now” — the fetch a chat assistant makes when a user asks about your product, or a browser-use agent completing a task. Training is “a crawler taking your content to train or fine-tune a model.”

For each behaviour you can now choose one of three states: block on all pages, block only on pages that display ads, or do not block at all. That granularity is genuinely new. Until now the practical choice was a single “block AI bots” toggle that lumped a model-training scraper together with the fetch that puts your brand into a ChatGPT answer.

Timeline of Cloudflare's 2026 AI crawler policy changeTimeline with four points: July 1, 2026, Cloudflare announces Search, Agent and Training crawler categories; July 14, 2026, analysts flag the multi-purpose crawler risk; August 10, 2026, today, with 36 days remaining to audit settings; and September 15, 2026, when the new defaults take effect for new domains and all free-tier sites.From announcement to enforcementJul 1, 2026Three categoriesannouncedJul 14, 2026Multi-purpose botrisk flaggedAug 10, 2026Today — 36 daysto audit settingsSep 15, 2026New defaultstake effectTodayEnforcement dateSource: Cloudflare announcement, July 1, 2026

The commercial goal is not subtle. Cloudflare is pairing the controls with a shift from its earlier Pay Per Crawl marketplace to a model it calls Pay Per Use, which pays publishers when their content shapes an AI answer rather than only when a page is fetched, with Ceramic.ai and You.com named as early partners. Alongside it the company shipped BotBase, a searchable directory of bot classifications for Enterprise customers, and extended its Content Signals robots.txt vocabulary with use=immediate, use=reference, and use=full preferences.

What Changes on September 15, 2026 — and Who It Applies To

The default configuration arriving on September 15 is narrower than the headlines suggest, and reading it precisely is the difference between a five-minute settings check and an unexplained traffic decline in October.

Cloudflare's three crawler categories and their September 15 defaultsThree columns. Search crawlers, which index content to answer questions later, are allowed by default. Agent crawlers, which act in real time on a person's behalf, are blocked by default on pages that display ads. Training crawlers, which collect content to train or fine-tune models, are also blocked by default on pages that display ads. Below, a note explains that a crawler performing more than one of these behaviours is governed by the most restrictive rule that applies, so Googlebot, Bingbot and Applebot are blocked wherever Training is blocked.Three behaviours, three controlsDefaults from September 15, 2026SEARCHIndexes your content toanswer questions laterAllowedon all pagesAGENTActs in real time ona person's behalfBlockedon ad-displaying pagesTRAININGCollects content to trainor fine-tune a modelBlockedon ad-displaying pagesThe catch: multi-purpose crawlersA crawler that performs more than one behaviour follows the most restrictive rule that applies to it.Googlebot, Bingbot and Applebot combine Search with Training — so they are blocked wherever Training is.Source: Cloudflare, July 1, 2026

The default itself

  • Search crawlers stay allowed. Nothing in the default blocks indexing behaviour on its own.
  • Training and Agent crawlers are blocked on ad-displaying pages only. Pages without ads are untouched by the default.
  • Multi-purpose crawlers follow the strictest applicable rule. This is the clause that does the damage. More on it below.

Who gets switched

The new defaults apply to brand-new domains connected to Cloudflare, to new sites added inside an existing account, and to every site on the free tier. If you are an existing customer who has already configured bot settings deliberately, your configuration is not overwritten — but that carve-out is thinner than it sounds. Plenty of sites have never touched the bot panel and are still sitting on Cloudflare's original defaults, and agencies routinely add client domains to existing accounts, which puts those new zones squarely in scope.

Pro Tip

Inventory before you audit. If you manage more than a handful of zones, list every domain added to your Cloudflare account since July 1 — those are the ones landing on the new defaults regardless of how the parent account is configured.

The Multi-Purpose Crawler Problem: How a Training Block Reaches Googlebot

Here is the sentence that should change what you do this month. Cloudflare blocks multi-purpose bots according to all of their behaviours. Googlebot, Bingbot, and Applebot do not fit neatly into one bucket: they index for search and feed AI systems. Under the new model, a crawler that performs both Search and Training is treated as a Training crawler everywhere Training is blocked.

Follow that through. A publisher who reasonably decides “I don't want my work training somebody's model” sets Training to blocked. Because Googlebot is a multi-purpose crawler, that setting also removes Googlebot's access — on the ad-carrying templates under the default, or sitewide if the block is set to all pages. The intent was to protect content from model training. The effect includes losing the crawler that feeds classic organic search.

Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge.

Matthew Prince, co-founder and CEO, Cloudflare

Prince has been explicit that the design is deliberate pressure: Cloudflare hopes the rules will “encourage mixed-use crawlers to separate out search from agent use and training.” That is a defensible objective for the web at large. It is also a bet placed with your crawl budget while the negotiation plays out. Cloudflare noted that Google currently has access to roughly twice as much information as other AI companies because of how it structures its crawler policies, and that more than half of AI crawler traffic is spent re-fetching pages that have not changed.

The analysts who covered the change flagged the same trap. As Arcalea put it in an analysis first published July 14 and updated August 4: “A crawler that performs both Search and Training is treated as a Training crawler wherever Training is blocked.” Their warning about how this surfaces is the important part — site owners are unlikely to connect a sudden search traffic decline to a crawler configuration decision made months earlier.

This failure mode is silent

Blocking a crawler does not produce an alert, a Search Console message, or a ranking penalty you can point at. It produces a slow decay in crawl requests, then stale content in the index, then decline. By the time the graph bends, the change that caused it is several weeks in the past and rarely the first suspect.

Why You Cannot Signal Your Way Out of This

The instinct of most SEOs facing a crawler-permissions question is to reach for a file: robots.txt, llms.txt, a Content Signals directive. For this particular problem that instinct is wrong, and the evidence is unusually clear.

Google's official guidance on AI features states plainly that “you don't need to create new machine readable files, AI text files, or markup to appear in these features,” and adds that there is no special schema.org structured data to add either. Gary Illyes has said Google does not support llms.txt and has no plans to. On Cloudflare's own vocabulary, John Mueller said the Content Signals robots.txt directive has “no effects whatsoever for any crawler or LLM,” and that as far as he knows no crawler or LLM uses it.

The adoption data points the same way. Ahrefs tracked 137,210 domains publishing a valid llms.txt file and found that 97% of them received zero requests for the file during May 2026. Of the small remainder that were fetched at all, most requests came from GPTBot and Claude Code rather than from anything connected to search ranking or citation.

No effects whatsoever for any crawler or LLM.

John Mueller, Search Advocate, Google, on Cloudflare's Content Signals directive

The practical conclusion is uncomfortable but simple. A declarative file is a request that crawlers are free to ignore, and by their own account they are ignoring these. Cloudflare's block is enforcement at the network edge — it returns a refusal rather than expressing a preference. That is precisely why it works, and precisely why misconfiguring it has consequences a stray robots.txt line never had. If you want to check what your own robots directives currently say before touching anything at the edge, our free Robots.txt Checker will parse the file and show which user-agents you are actually addressing.

How to Audit Your Cloudflare AI Bot Settings Before September 15

This is a configuration review, not a strategy project. For most sites it is under fifteen minutes. Run it once per zone.

Step 1: Open the AI bot policy panel

In the Cloudflare dashboard, select the zone and navigate to Security → Settings → Configure AI bot policies. You are looking for the three named categories. If you only see a legacy single toggle, that zone has not been migrated yet and will inherit the new behaviour — note it and come back after September 15 to confirm.

Step 2: Set each category deliberately, not as a group

Decide Search, Agent, and Training separately. For most commercial sites that are not funded by display advertising, the defensible configuration is Search allowed, Agent allowed, and Training set according to your view on model training. If you do block Training, do it knowing it reaches Googlebot on the pages in scope — scope it as narrowly as you can rather than applying it sitewide.

Step 3: Identify which of your pages count as “ad-displaying”

The default block is scoped to pages that display ads, so the size of your exposure depends entirely on how much of your site carries ad units. A SaaS blog with no ads has almost no surface area here. A content site running display inventory across every article template has essentially its whole library in scope. Map this before you decide anything, because it converts an abstract policy into a concrete page count.

Step 4: Verify what crawlers actually receive

Settings pages describe intent; responses are the truth. Check the status codes and headers your monetised templates return to a Googlebot user-agent, and watch for 403s appearing where you expect 200s. Our free HTTP Header Checker shows the exact response headers a given URL returns, which is the quickest way to catch an edge rule doing something you did not intend.

Step 5: Set a baseline you can compare against in October

Record your current crawl stats from Search Console — total crawl requests, average response time, and newly-discovered URLs — before September 15. If something does break, crawl data moves weeks before rankings do, and having a pre-change baseline turns a vague suspicion into a diagnosis. A full Technical SEO Audit gives you a dated snapshot of crawlability and indexation to compare against afterwards.

Pro Tip

Diary the check. Put a calendar reminder for September 16 and another for October 15. The first confirms the defaults landed the way you expected; the second is early enough to catch a crawl decline while it is still cheap to reverse.

Should You Block Training Crawlers at All?

The honest answer is that it depends on a single question: does someone need to land on your page for you to make money from it?

If ad impressions are the revenue, an AI system that consumes your article and answers the question without sending the visit is taking the product and leaving nothing. Blocking is a rational commercial position, and it is exactly the case Cloudflare designed the ad-page scoping around. The pushback on the open web — that AI Overviews cut outbound clicks by roughly 40% in one field study — lands hardest on precisely these publishers.

If your content exists to generate leads, trials, or sign-ups, the calculation inverts. Guidance published alongside the change put it bluntly: the only time blocking makes sense is “if you're legitimately making money from the content that's there… if you're actually needing people to visit the page for ad revenue in order to make money.” Service businesses, portfolios, and product-supporting blogs get more from being discoverable than from being protected. Being absent from training data and from agent fetches does not protect that kind of business — it makes it harder to surface in the answers where buyers now start.

There is a middle path worth naming. Blocking Training while leaving Agent allowed keeps you out of the next model's weights but still lets a live assistant fetch your page when a user asks about you — the fetch most likely to produce a citation and a click today. Given that Agent and Training carry the same default, this is a deliberate choice you have to make rather than one you inherit. If you want to see how often your pages currently surface in generative results before you change anything, our AI Overview Analyzer is a reasonable place to establish that baseline.

Whichever way you go, make it a decision. The failure mode worth avoiding is not choosing wrong — it is inheriting a default that does not match your business model and finding out in November.

Tools to Help You Audit Crawl Access

Most of this work is verification: confirming that what your edge configuration says matches what crawlers actually receive. These are free and take a few minutes each.

What to Expect Next

The obvious thing to watch is whether the pressure works. Cloudflare's stated aim is to push the major crawlers into separating search from training, and the cleanest resolution for everyone would be Google shipping a distinct user-agent for each behaviour, so that blocking one no longer implies blocking the other. If that happens in the next few quarters, the trap described in this article closes on its own.

Until then, expect the workaround economy. Google and Apple already offer separate opt-out mechanisms that sit outside Cloudflare's classification, and regulators are moving on parallel tracks — the UK has been pushing its own opt-out requirements, and litigation over training data continues on both sides of the Atlantic. A negotiated settlement between infrastructure providers, AI companies, and publishers is a multi-year process, not a September one.

The specific thing to watch on your own properties is simpler: crawl request volume in Search Console during the last two weeks of September. If Cloudflare's classification catches a crawler you meant to allow, that number moves first. Keep an eye on it before you start looking for algorithmic explanations — we have written before about how easily a crawl-side change gets misread as a ranking event.

Frequently Asked Questions

Key Takeaways

Cloudflare's change is a reasonable answer to a real problem, and the three-category model is a genuine improvement over a single on/off switch. The risk is not the policy — it is that the most restrictive-rule clause converts a content-licensing decision into a search-indexing decision without saying so out loud. Thirty-six days is plenty of time to check, and almost nobody will.

Your Action Plan:

  • Open Security → Settings → Configure AI bot policies on every zone you manage, before September 15.
  • Set Search, Agent, and Training separately — and assume any Training block also applies to Googlebot on the pages in scope.
  • Map which of your templates carry ad units; that is the true scope of the default block.
  • Record Search Console crawl stats now as a baseline, and re-check them in late September.
  • Do not rely on llms.txt or Content Signals to express any of this — Google has said neither does anything.

If you are working through how AI crawling and citation fit together more broadly, our breakdown of how citations differ by engine and our look at what llms.txt actually does both go deeper, or browse the full articles library for ongoing coverage.

Related Free SEO Tools

More Articles