Guides · CDN

Why Cloudflare blocks Googlebot, and how to check whether it is happening to you

9 min read · published 30 September 2026

A site can have a perfect robots.txt and still be invisible to Google, because the CDN in front of it answers crawlers with a challenge page. Cloudflare is the CDN most sites use, so it is the one most often blamed, but the same mechanisms exist at every bot-protection vendor. This guide explains how it happens and how to prove or rule it out in ten minutes.

The mechanisms

How to see it

In Cloudflare. Open Security, then Events (called Security Analytics on newer dashboards), and filter by the user agent of the crawler or by the “verified bot” field. Every challenge, block or rate limit is listed with the rule that caused it. If Googlebot appears with action “Managed Challenge” or “Block”, you have found it.

In Search Console. Settings, then Crawl stats, shows Googlebot’s own view: the host status section flags robots.txt fetch failures and server connectivity problems, and the response breakdown shows a spike in 403s or “other client error” when a challenge went up. The URL Inspection tool’s live test fetches a page as Googlebot right now and shows the HTTP status and the rendered result.

With the checker. Run the canaicrawl checker on the page. It first requests it as a normal browser. If that request is refused with a 403 or a challenge, the report says “our checker was refused” and marks every bot unverified rather than blocked, because a 403 to us proves nothing about the real Googlebot, whose requests come from Google’s verified ranges. If the browser request succeeds and a specific bot gets a 403, that bot is being refused on its user agent alone, which points at a WAF rule or the AI-bot toggle. Either way the report shows which CDN answered and whether it served a challenge page.

In your logs, if you have them: look for the crawler’s user agent and whether the requests reach the origin at all. Requests that never arrive were stopped at the edge.

How to fix it without turning protection off

The same story at other CDNs

Akamai Bot Manager, Imperva, AWS WAF Bot Control, Fastly’s Next-Gen WAF, Vercel’s bot protection and the hosting-provider firewalls all have a “known good bots” list and a way to challenge everything else. The failure modes are identical: a challenge or rate limit applied above the bot allow-list, or a rule keyed on a string that search crawlers also carry. The checker names the CDN it detected in the response so you know which dashboard to open.

Check it on your site

More guides: How to block AI crawlers · robots.txt for AI crawlers · What is llms.txt · noindex vs robots.txt · Content-Signal explained.

Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.