Bytespider
ByteDance’s crawler, believed to collect training data for its models. It has no public documentation and is often reported to ignore robots.txt.
- Operator
- ByteDance
- Purpose
- Model training. Crawlers that collect pages to train models. Blocking them does not affect search or AI answers.
- robots.txt token
Bytespider- User agent
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)- Documentation
- None published by the operator.
Block Bytespider
Add a group naming it to /robots.txt. Compliant crawlers stop before requesting any page. A firewall rule on the user agent is the only way to enforce it against crawlers that ignore robots.txt.
User-agent: Bytespider
Disallow: /
Allow Bytespider when everything else is blocked
A group naming the bot wins over User-agent: *, so an explicit allow lets it in while your catch-all rules stay in place.
User-agent: Bytespider
Allow: /
User-agent: *
Disallow: /
Does Bytespider reach your site?
The checker applies your robots.txt the way Bytespider does, then requests your page with the user agent above, alongside 29 other bots.
Other ai training: GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, GoogleOther, Meta-ExternalAgent.
Build a complete file with the robots.txt generator.