Guides · Blocking

How to block AI crawlers: robots.txt, firewall rules and what each one costs you

9 min read · published 30 September 2026

Blocking AI crawlers is a three-layer job: tell them to stay away, refuse the ones that don’t listen, and check that you haven’t locked out the bots you depend on. This guide walks through all three and the trade-offs behind each choice.

First decide what you are actually blocking

“AI crawlers” is four different things, and the cost of blocking each one is different.

The robots.txt generator has presets for each of these lines: block training only, block all AI, or block everything but search engines.

Layer 1: robots.txt

robots.txt is a request. Every major operator documents that its crawlers read it, and the training crawlers from OpenAI, Anthropic, Google, Apple, Common Crawl and Meta all say they honour a Disallow. One group per bot, or several bots in one group:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Bytespider
Disallow: /

User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

User-agent: *
Allow: /

Three details trip people up. A group naming a bot replaces the User-agent: * group for that bot rather than adding to it, so a bot with its own group ignores your catch-all rules entirely. Matching is on the product token, case-insensitively, so gptbot and GPTBot are the same. And the file must be served as plain text at /robots.txt on the exact host the pages live on; a robots.txt on the apex domain does nothing for www. The robots.txt guide covers the full syntax.

Layer 2: refuse at the server or CDN

A crawler that ignores robots.txt, or a scraper pretending to be one, only stops when the server refuses it. Rules keyed on the user-agent string are the simplest enforcement.

nginx, in the server block:

if ($http_user_agent ~* "(GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent)") {
    return 403;
}

Apache, in .htaccess or the virtual host:

RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent) [NC]
RewriteRule .* - [F,L]

Cloudflare has a one-click “block AI bots” setting under Security that refuses the AI crawlers on its verified-bot list, and since 2025 new zones get it on by default. For finer control, a WAF custom rule can match the user agent or, for verified bots, the bot category. Vercel, Akamai, Imperva and AWS WAF have equivalent bot-management rules. Whatever the tool, keep /robots.txt itself reachable by everyone: a crawler that cannot read robots.txt cannot learn that you want it gone.

User-agent rules stop honest bots and lazy scrapers. A scraper that changes its user agent to look like Chrome walks past them, which is why the operators that publish their IP ranges (OpenAI, Google, Microsoft and Apple do) are the ones you can verify at the network layer.

Layer 3: don’t block what you need

Most accidental damage comes from rules that were meant for AI bots and caught search engines too. The usual causes:

Verify it, then keep verifying

Run the checker on your home page and on one deep page. It applies your robots.txt the way each crawler does and shows the matching line, then requests the page with each bot’s user agent so you see what the server actually answers. A bot you meant to block should show “Blocked · robots.txt” or a 403; a bot you rely on should show 200. Pick the “search and AI answers, no training” goal if that is your policy, so deliberate blocks stop counting against you.

Then check your access logs for the tokens you blocked. A compliant crawler disappears within days. One that keeps coming is ignoring robots.txt, and only your server rule stops it.

Legal signals, briefly

Two machine-readable rights reservations exist alongside robots.txt: the TDM-Reservation: 1 HTTP header from the W3C TDMRep community group, which reserves text-and-data-mining rights under the EU copyright directive’s opt-out, and Cloudflare’s Content-Signal lines in robots.txt. Neither blocks anything by itself. They state your policy in a form a court or a compliant operator can read. The Content-Signal guide and the noindex guide explain both.

Check it on your site

More guides: robots.txt for AI crawlers · What is llms.txt · noindex vs robots.txt · Content-Signal explained · Cloudflare blocking Googlebot.

Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.