GPTBot
Collects pages that may be used to train OpenAI’s models. Blocking it does not affect ChatGPT search, which uses OAI-SearchBot.
- Operator
- OpenAI
- Purpose
- Model training. Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
- robots.txt token
GPTBot- User agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.1; +https://openai.com/gptbot- Documentation
- platform.openai.com/docs/bots
Block GPTBot
Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.
User-agent: GPTBot
Disallow: /
Allow GPTBot when everything else is blocked
A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.
User-agent: GPTBot
Allow: /
User-agent: *
Disallow: /
Does GPTBot reach your site?
The checker applies your robots.txt the way GPTBot does, then requests your page with the user agent above, alongside 29 other bots.
Also run by OpenAI: OAI-SearchBot, ChatGPT-User.
Other ai training: ClaudeBot, Google-Extended, CCBot, Applebot-Extended, GoogleOther, Meta-ExternalAgent, Bytespider.
Build a complete file with the robots.txt generator.