Guides · Blocking
How to block AI crawlers: robots.txt, firewall rules and what each one costs you
Blocking AI crawlers is a three-layer job: tell them to stay away, refuse the ones that don’t listen, and check that you haven’t locked out the bots you depend on. This guide walks through all three and the trade-offs behind each choice.
First decide what you are actually blocking
“AI crawlers” is four different things, and the cost of blocking each one is different.
- Training crawlers such as GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent and Bytespider collect pages to train models. Blocking them costs you nothing visible today: no search ranking, no AI answer citations. This is the group most sites choose to block.
- AI search crawlers such as OAI-SearchBot, PerplexityBot and Claude-SearchBot build the indexes that ChatGPT search, Perplexity and Claude answer from and link to. Blocking them removes you from those answers, which is a real and growing traffic source.
- User-triggered fetchers such as ChatGPT-User, Claude-User and Perplexity-User load a page because a person asked an assistant about it. Blocking them means the assistant tells that person it cannot open your page. Several operators state these fetchers do not honour robots.txt because a human requested the page.
- Robots.txt tokens such as Google-Extended and Applebot-Extended are not crawlers at all. They tell Google and Apple whether pages their normal crawlers already fetched may be used for AI. Disallowing them does not touch Google Search or Siri.
The robots.txt generator has presets for each of these lines: block training only, block all AI, or block everything but search engines.
Layer 1: robots.txt
robots.txt is a request. Every major operator documents that its crawlers read it, and the training crawlers from OpenAI, Anthropic, Google, Apple, Common Crawl and Meta all say they honour a Disallow. One group per bot, or several bots in one group:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Bytespider
Disallow: /
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: *
Allow: /
Three details trip people up. A group naming a bot replaces the User-agent: * group for that bot rather than adding to it, so a bot with its own group ignores your catch-all rules entirely. Matching is on the product token, case-insensitively, so gptbot and GPTBot are the same. And the file must be served as plain text at /robots.txt on the exact host the pages live on; a robots.txt on the apex domain does nothing for www. The robots.txt guide covers the full syntax.
Layer 2: refuse at the server or CDN
A crawler that ignores robots.txt, or a scraper pretending to be one, only stops when the server refuses it. Rules keyed on the user-agent string are the simplest enforcement.
nginx, in the server block:
if ($http_user_agent ~* "(GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent)") {
return 403;
}
Apache, in .htaccess or the virtual host:
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent) [NC]
RewriteRule .* - [F,L]
Cloudflare has a one-click “block AI bots” setting under Security that refuses the AI crawlers on its verified-bot list, and since 2025 new zones get it on by default. For finer control, a WAF custom rule can match the user agent or, for verified bots, the bot category. Vercel, Akamai, Imperva and AWS WAF have equivalent bot-management rules. Whatever the tool, keep /robots.txt itself reachable by everyone: a crawler that cannot read robots.txt cannot learn that you want it gone.
User-agent rules stop honest bots and lazy scrapers. A scraper that changes its user agent to look like Chrome walks past them, which is why the operators that publish their IP ranges (OpenAI, Google, Microsoft and Apple do) are the ones you can verify at the network layer.
Layer 3: don’t block what you need
Most accidental damage comes from rules that were meant for AI bots and caught search engines too. The usual causes:
- A
User-agent: *group withDisallow: /and no explicit group for Googlebot or Bingbot. - A firewall rule on the word
bot, which matches Googlebot, bingbot and Applebot. - Cloudflare’s Bot Fight Mode or a managed challenge on all traffic, which crawlers cannot solve. The Cloudflare guide is about exactly this.
- Blocking OAI-SearchBot while meaning to block GPTBot. They are different bots with different jobs.
Verify it, then keep verifying
Run the checker on your home page and on one deep page. It applies your robots.txt the way each crawler does and shows the matching line, then requests the page with each bot’s user agent so you see what the server actually answers. A bot you meant to block should show “Blocked · robots.txt” or a 403; a bot you rely on should show 200. Pick the “search and AI answers, no training” goal if that is your policy, so deliberate blocks stop counting against you.
Then check your access logs for the tokens you blocked. A compliant crawler disappears within days. One that keeps coming is ignoring robots.txt, and only your server rule stops it.
Legal signals, briefly
Two machine-readable rights reservations exist alongside robots.txt: the TDM-Reservation: 1 HTTP header from the W3C TDMRep community group, which reserves text-and-data-mining rights under the EU copyright directive’s opt-out, and Cloudflare’s Content-Signal lines in robots.txt. Neither blocks anything by itself. They state your policy in a form a court or a compliant operator can read. The Content-Signal guide and the noindex guide explain both.
Check it on your site
More guides: robots.txt for AI crawlers · What is llms.txt · noindex vs robots.txt · Content-Signal explained · Cloudflare blocking Googlebot.
Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.