
Until recently, deciding whether to let artificial intelligence crawlers read your website was a philosophical question. As of September 2026 it is a configuration setting with a default — and for new domains on Cloudflare, that default now blocks two of the three categories.
This guide covers what the three categories are, what changed, and how to decide which ones to allow. The right answer is different for a news publisher, an ecommerce store, and a business whose website is mostly a brochure.
The Three Kinds of AI Bot
Not all automated traffic from artificial intelligence companies does the same job. Cloudflare separates it into three categories, and that distinction is the whole decision.
Search crawlers index your pages so they can be cited and linked in generated answers. This is the closest analogue to traditional search indexing — the bot reads your content, and in return your site can appear as a source with a link back.
Agent bots fetch pages on behalf of a person who asked a question right now. Someone asks an assistant to compare three suppliers; the assistant visits all three sites and reads them. This is a real user with a machine intermediary.
Training crawlers collect content to train future models. Your content becomes part of the model’s weights. There is no link back, no attribution, and no traffic.
Most site owners who say they want to “block AI” mean Training. Blocking all three is a different decision with different consequences.
What Changed in September 2026
From 15 September 2026, new domains onboarding to Cloudflare have the Training and Agent categories blocked by default on pages that display advertising. Search remains allowed by default.
Existing domains keep whatever settings they already had — the change applies to new onboarding. But the shift in default signals where the industry is heading: content owners are being handed the presumption of control rather than having to opt out.
The controls themselves are available on every Cloudflare plan, including the free tier.
Should You Block Them?
The answer depends on what your website is for.
If your site earns money from traffic — advertising, affiliate, subscription — Training crawlers take content and return nothing. Blocking them is straightforward. Agent bots are more nuanced: they represent real users, but users who may never visit your page directly.
If your site is a business brochure — the goal is being found and contacted. Search crawlers help; being cited in a generated answer with a link is closer to free advertising than to theft. Blocking Search would be self-harm. Training is a judgement call with little practical downside either way.
If your site is ecommerce — Agent traffic matters more than either of the others. Assistants that compare products on behalf of shoppers are a growing acquisition channel. Blocking Agent bots removes your products from those comparisons.
If your content is your product — original research, journalism, paid courses — Training blocking is the entire point, and Agent needs a deliberate decision about whether you want your material summarised without a visit.
How to Configure the Controls
The settings live under the Bot Management section of the Cloudflare dashboard.
- Review the current setting for each category — Search, Agent, Training — before changing anything.
- Set each category deliberately rather than applying one blanket rule across all three.
- Check analytics after a fortnight: Cloudflare reports which specific bots hit the site and how often.
- Pair with a robots.txt policy for crawlers that respect it, understanding that network-layer blocking is what actually enforces the decision.
For sites already running custom firewall rules, the crawler controls sit alongside them rather than replacing them — see How Do You Configure Cloudflare WAF Rules? for how the two layers interact.
Frequently Asked Questions
Will blocking AI crawlers hurt my search rankings?
No. Conventional search engine crawlers such as Googlebot and Bingbot are a separate category and are not affected by these controls. Blocking Training or Agent bots does not change how traditional search indexes your site.
Do AI crawlers respect robots.txt?
Reputable ones do. Others ignore it, and some disguise themselves as ordinary browser traffic. This is why network-layer blocking exists — it enforces the decision rather than politely requesting compliance.
Can I allow some companies and block others?
Yes. The category controls are the blunt instrument; Cloudflare also exposes per-bot visibility so individual crawlers can be allowed or blocked by name. Most businesses start with categories and refine later if a specific bot becomes a problem.
Set Your Crawler Policy Deliberately
Artificial intelligence crawler policy is now a live configuration decision rather than a theoretical one, and the right answer depends on how your website actually earns its keep. ANP Technology helps businesses review their current bot posture, set category controls deliberately, and measure what the change does to traffic.
Talk to ANP Technology about your Cloudflare configuration →






