Guides · CDN
Why Cloudflare blocks Googlebot, and how to check whether it is happening to you
A site can have a perfect robots.txt and still be invisible to Google, because the CDN in front of it answers crawlers with a challenge page. Cloudflare is the CDN most sites use, so it is the one most often blamed, but the same mechanisms exist at every bot-protection vendor. This guide explains how it happens and how to prove or rule it out in ten minutes.
The mechanisms
- Bot Fight Mode and Super Bot Fight Mode classify traffic as verified bot, likely automated or definitely automated. Verified bots, which include Googlebot, Bingbot and the major AI crawlers, are supposed to pass. The trouble starts when a request is not recognised as verified: an IP outside the published ranges, a new crawler, or a rule that runs before verification.
- Managed challenges and “I’m Under Attack” mode serve a JavaScript or interactive challenge to every visitor. Crawlers do not execute challenge scripts, so they receive a 403 with a
cf-mitigated: challengeheader and index nothing. A challenge left on after an attack is the single most common cause of a site vanishing from search. - WAF custom rules written for humans catch crawlers by accident: a block on the word
botin the user agent, a country block that includes the crawler’s data centre, a rate limit that a crawl exceeds, or a block on hosting-provider networks. - The “block AI bots” toggle refuses the AI crawlers on Cloudflare’s verified-bot list. It is on by default for zones created since mid-2025. That is what many owners want, but it also blocks OAI-SearchBot, PerplexityBot and Claude-SearchBot, which put you in AI answers, and some owners do not realise the toggle covers those.
- Firewall rules on the origin behind Cloudflare, which see every request coming from Cloudflare’s IPs and sometimes rate-limit or block them wholesale.
How to see it
In Cloudflare. Open Security, then Events (called Security Analytics on newer dashboards), and filter by the user agent of the crawler or by the “verified bot” field. Every challenge, block or rate limit is listed with the rule that caused it. If Googlebot appears with action “Managed Challenge” or “Block”, you have found it.
In Search Console. Settings, then Crawl stats, shows Googlebot’s own view: the host status section flags robots.txt fetch failures and server connectivity problems, and the response breakdown shows a spike in 403s or “other client error” when a challenge went up. The URL Inspection tool’s live test fetches a page as Googlebot right now and shows the HTTP status and the rendered result.
With the checker. Run the canaicrawl checker on the page. It first requests it as a normal browser. If that request is refused with a 403 or a challenge, the report says “our checker was refused” and marks every bot unverified rather than blocked, because a 403 to us proves nothing about the real Googlebot, whose requests come from Google’s verified ranges. If the browser request succeeds and a specific bot gets a 403, that bot is being refused on its user agent alone, which points at a WAF rule or the AI-bot toggle. Either way the report shows which CDN answered and whether it served a challenge page.
In your logs, if you have them: look for the crawler’s user agent and whether the requests reach the origin at all. Requests that never arrive were stopped at the edge.
How to fix it without turning protection off
- Exempt verified bots. Add a WAF custom rule that skips the challenge and bot-management rules when the request is from a verified bot (the
cf.client.botfield), or, more narrowly, when the verified-bot category is “Search Engine Crawler”. Put it above the rules that challenge. - Turn the challenge off, or scope it. Under Attack mode is for attacks. When one ends, switch back and, if you must keep a challenge, apply it to the paths that were attacked rather than to the whole site.
- Check the AI-bot toggle against your policy. If you want AI search traffic, either turn it off and block training crawlers in robots.txt and a WAF rule instead, or use the per-bot controls newer dashboards offer.
- Never challenge robots.txt. Add a rule that always allows
/robots.txt,/sitemap.xmland/llms.txt. A crawler that cannot read robots.txt treats the site as fully disallowed. - Re-check. Run the checker again, and use Search Console’s live test to confirm Googlebot itself gets a 200. Changes at the edge apply within seconds; Search Console’s crawl stats lag by a day or two.
The same story at other CDNs
Akamai Bot Manager, Imperva, AWS WAF Bot Control, Fastly’s Next-Gen WAF, Vercel’s bot protection and the hosting-provider firewalls all have a “known good bots” list and a way to challenge everything else. The failure modes are identical: a challenge or rate limit applied above the bot allow-list, or a rule keyed on a string that search crawlers also carry. The checker names the CDN it detected in the response so you know which dashboard to open.
Check it on your site
More guides: How to block AI crawlers · robots.txt for AI crawlers · What is llms.txt · noindex vs robots.txt · Content-Signal explained.
Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.