Bot directory · AI training

AI training · Common Crawl

CCBot

Builds the free Common Crawl archive, which many AI models have been trained on.

Operator
Common Crawl
Purpose
Model training. Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
robots.txt token
CCBot
User agent
CCBot/2.0 (https://commoncrawl.org/faq/)
Documentation
commoncrawl.org/ccbot

Block CCBot

Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.

User-agent: CCBot
Disallow: /

Allow CCBot when everything else is blocked

A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.

User-agent: CCBot
Allow: /

User-agent: *
Disallow: /

Does CCBot reach your site?

The checker applies your robots.txt the way CCBot does, then requests your page with the user agent above, alongside 29 other bots.

Other ai training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, GoogleOther, Meta-ExternalAgent, Bytespider.

Build a complete file with the robots.txt generator.