The Newsletter · English Edition

Your robots.txt says yes. Cloudflare says no.

On 15 September 2026, one Cloudflare setting changes. If your site sits behind Cloudflare, it isn't only your presence in ChatGPT that's at stake; it's Googlebot, on part of your pages. Here is the check to run, in plain English. Ten minutes, free, no technical skills required.

Originally published in French. This edition first appeared in Visible sur l'IA, my French-language newsletter on LinkedIn.

"How do I know if my site is blocked for AI?"

I get that question every week now. It's a good question, and the answer is rarely where people look. Most site owners check their robots.txt, see that everything is allowed, and conclude they're fine.

They're often not. And on 15 September 2026, the gap between what your robots.txt says and what actually happens gets wider.

10 min
what the whole check takes, start to finish
3
bot families Cloudflare starts sorting on 15 September
$0
AI Crawl Control is on every plan, free included

First: what is Cloudflare?

It's a concierge posted at the front of your building. Your site is the building, and everyone who wants in, visitors and bots alike, walks past him first. He speeds up your pages, he protects you from attacks, and he decides who gets in.

There's a good chance this is your situation without you ever having decided it: usually it's your host or your agency who set it up. To find out in ten seconds, type your address followed by /cdn-cgi/trace in your browser. Example: yoursite.com/cdn-cgi/trace. If a technical page appears, you're on Cloudflare. If you get an error, you're not, and you can go about your day.

This is also why robots.txt alone tells you nothing. Robots.txt is a sign on the door. Cloudflare is the bouncer, and he stands in front of the sign. A bot that gets turned away at the edge never reaches your file, and never shows up in your server logs either.

Diagram: AI crawlers such as GPTBot, ClaudeBot, PerplexityBot and CCBot are stopped by Cloudflare before they ever reach the site's robots.txt file
Blocking happens at the edge, before your robots.txt is ever read. Nothing shows up in your logs.

What changes on 15 September

Cloudflare is going to sort bots into three families: the ones that index for search, the ones that train AI, and autonomous agents.

Two things change, and they should not be confused.

The first is the default setting: search allowed, training and agents blocked on pages that display advertising. That one will apply on its own only to new domains, to new sites added by existing customers, and to free accounts that have never changed anything. If you pay for a plan and someone has already configured things on your end, your settings won't move.

The second concerns everyone, and this is the real trap. Plenty of bots do two jobs at once. Googlebot indexes for Google and also feeds AI. Same for Bingbot and Applebot. To those bots, Cloudflare will apply the strictest instruction you have given.

Table of the three Cloudflare bot families from 15 September 2026: search, training, agent, plus mixed bots such as Googlebot, Bingbot and Applebot
Three families, three consequences. The fourth row is the one that can cost you Google.

Translation: if your concierge has been told to block AI training, he will also block Googlebot on your pages that display advertising. It doesn't matter how long you've been with Cloudflare, it doesn't matter which plan you're on, including if someone ticked the old "Block AI bots" box a year ago and nobody has touched it since.

The most ordinary case, and it's the one I see everywhere: a site blocked GPTBot and CCBot a year ago, thinking it was protecting itself from AI training only. Along the way, someone also ticked "Block AI bots". On 15 September, that site loses Googlebot on every page that displays advertising.

At that point it's no longer about being cited or not by ChatGPT. It's about dropping out of plain old Google indexing, the one that has been bringing in traffic for fifteen years, on part of its pages.

The setting you thought was protecting you can make you disappear from Google. That is exactly why this gets checked now, and not on the 16th.

The check, step by step

1

Log in to Cloudflare

Go to dash.cloudflare.com and log in. You don't have the credentials? That's common. Ask your agency or your host, nine times out of ten they're the ones holding them.

2

Open your domain

Click your domain name in the list. If several domains are in the account, do this for each one that matters to you.

3

Find "Block AI bots"

Open the security settings, filter on bot traffic, and find Block AI bots. Just look at which position it's in: block everything, block only on pages with advertising, or block nothing.

That single toggle is what turns into the strictest-rule instruction on 15 September. It's worth thirty seconds of your attention.

4

Go bot by bot in AI Crawl Control

To go further, open AI Crawl Control. You'll see the list of bots, one by one, with who gets through and who gets turned away. GPTBot, ClaudeBot, PerplexityBot, and the rest.

This is where the decision actually belongs: this one I want reading me, that one I don't.

And what does it cost?

Nothing. AI Crawl Control is available on every Cloudflare plan, free plan included. Nothing to buy, nothing to install. Just look, and decide.

Two questions I get all the time

"I don't have a Cloudflare account, so this doesn't concern me, right?"

It does. And this is the point I haven't pushed hard enough so far. You don't need an account to be affected: your site's traffic only has to pass through Cloudflare. Very often it's your host, your agency or the previous developer who switched it on, sometimes bundled into the offer by default. The account exists. It's simply not in your name, and the settings apply to your site all the same.

"And how do I check whether I'm affected, without creating an account?"

From the outside, free, with nothing to install.

First the /cdn-cgi/trace test, to find out whether you're going through Cloudflare. Then, after 15 September, open Google Search Console, take one of your pages that displays advertising, run URL Inspection and then Test live URL. If Googlebot can no longer fetch the page, you have your answer in thirty seconds, without opening a single Cloudflare account.

And if the test comes back bad, you won't be able to fix it yourself: you'll need to ask your host or your agency for access. You may as well write to them now rather than on the 16th.

The one line to remember

Because that's really what this is about. I'm not telling you to throw the door wide open to AI. Refusing training is a perfectly defensible choice, especially if your content is how you make a living. What I am telling you is that it has to be your choice, made knowingly, and not an instruction the concierge took from someone else.

Ten minutes this week. You'll finally know who you're letting in.


Scope and sources

The new default (search allowed, training and agents blocked on pages with ads) applies automatically only to new domains, to new sites added by existing customers, and to free accounts that have never been modified. The strictest-rule logic on mixed crawlers applies to anyone already blocking training, whatever the plan.

Loss of Googlebot concerns pages that display advertising, not the whole site.

Sources: Cloudflare developer documentation (Block AI bots, AI Crawl Control) and the Cloudflare blog, "Your site, your rules". Checked in August 2026.

En français ?

This newsletter lives in French

Visible sur l'IA is my French-language newsletter on LinkedIn: studies, practical SEO and AI-visibility decoding, every month. If you read French, that's where each edition lands first.

Researched, tested and written by Roxane Pinault.