Bot directory · AI training

AI training · Meta

Meta-ExternalAgent

Collects pages to train Meta’s AI models and to improve its products.

Operator
Meta
Purpose
Model training. Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
robots.txt token
Meta-ExternalAgent
User agent
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
Documentation
developers.facebook.com/docs/sharing/webmasters/web-crawlers

Block Meta-ExternalAgent

Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.

User-agent: Meta-ExternalAgent
Disallow: /

Allow Meta-ExternalAgent when everything else is blocked

A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.

User-agent: Meta-ExternalAgent
Allow: /

User-agent: *
Disallow: /

Does Meta-ExternalAgent reach your site?

The checker applies your robots.txt the way Meta-ExternalAgent does, then requests your page with the user agent above, alongside 29 other bots.

Also run by Meta: Meta-ExternalFetcher.

Other ai training: GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, GoogleOther, Bytespider.

Build a complete file with the robots.txt generator.