Skip to content
Training Crawler

Bytespider

Operated by ByteDance · Widely attributed to ByteDance · Purpose not officially documented

GEO Recommendation: Optional to Block

Bytespider is widely reported to be ByteDance's web crawler used for AI training data. Unlike every other entry in this directory, ByteDance has not published official crawler documentation confirming its purpose, robots.txt behavior, or IP ranges — treat the details below as third-party-reported, not vendor-confirmed.

Owner
ByteDance
Trigger
Automated crawl (reported)
Used for AI Training
Reported, unconfirmed by ByteDance
robots.txt control
Inconsistent, per third-party reports
Impact if blocked
A robots.txt rule may not reliably stop this crawler, per widespread third-party reports — server-level blocking is commonly recommended instead
How to identify
User-Agent: Bytespider

What it does

Bytespider is widely reported across the industry as ByteDance's crawler for gathering AI training data. However, ByteDance has not published a vendor documentation page confirming this purpose, its robots.txt behavior, or its published IP ranges.

Multiple independent third-party sources describe Bytespider as high-volume and, at times, not reliably respecting robots.txt directives — but because this comes from third-party observation rather than ByteDance's own documentation, ELNIQ presents it as reported, not confirmed.

Why this is optional to block

  • No official ByteDance documentation confirms this crawler's purpose or behavior
  • Multiple independent reports describe it as not always respecting robots.txt
  • Consider server- or WAF-level blocking if reliable enforcement matters to you

Identification & robots.txt

User-Agent identifier
Bytespider
Verification

Because ByteDance has not published official documentation, ELNIQ cannot verify this crawler's behavior against a primary source. Treat identification as best-effort, based on the User-Agent string alone.

Robots.txt guidance — to opt out
User-agent: Bytespider Disallow: /