Bot directory · AI training

AI training · ByteDance

Bytespider

ByteDance’s crawler, believed to collect training data for its models. It has no public documentation and is often reported to ignore robots.txt.

Operator
ByteDance
Purpose
Model training. Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
robots.txt token
Bytespider
User agent
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)
Documentation
None published by the operator.

Block Bytespider

Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.

User-agent: Bytespider
Disallow: /

Allow Bytespider when everything else is blocked

A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.

User-agent: Bytespider
Allow: /

User-agent: *
Disallow: /

Does Bytespider reach your site?

The checker applies your robots.txt the way Bytespider does, then requests your page with the user agent above, alongside 29 other bots.

Other ai training: GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, GoogleOther, Meta-ExternalAgent.

Build a complete file with the robots.txt generator.