Bot directory · AI training

AI training · Google

Google-Extended

Not a crawler. A robots.txt token that tells Google whether pages Googlebot fetched may train Gemini and be used for grounding. Disallowing it does not affect Google Search.

Operator
Google
Purpose
Model training. Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
robots.txt token
Google-Extended
User agent
None. This is a robots.txt token, not a crawler: the operator’s normal crawler fetches the page and this name only governs how it may be used.
Documentation
developers.google.com/search/docs/crawling-indexing/google-common-crawlers#google-extended

Block Google-Extended

Add a group naming it to /robots.txt. Nothing stops fetching; the operator’s crawler still reads the page for its normal purpose, but may not use it this way.

User-agent: Google-Extended
Disallow: /

Allow Google-Extended when everything else is blocked

A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.

User-agent: Google-Extended
Allow: /

User-agent: *
Disallow: /

Does Google-Extended reach your site?

The checker applies your robots.txt the way Google-Extended does, alongside 29 other bots.

Also run by Google: Googlebot, GoogleOther.

Other ai training: GPTBot, ClaudeBot, CCBot, Applebot-Extended, GoogleOther, Meta-ExternalAgent, Bytespider.

Build a complete file with the robots.txt generator.