Skip to content
Dataset Crawler

CCBot

Operated by Common Crawl · Feeds a shared, multi-company dataset

GEO Recommendation: Optional to Block

CCBot is the crawler for Common Crawl, a nonprofit that publishes an open web archive. Unlike single-company bots, this archive is downloaded and used for training by many different AI labs — blocking CCBot is a general opt-out from that shared dataset, not from any one company's models.

Owner
Common Crawl
Trigger
Automated crawl
Used for AI Training
Indirectly — Common Crawl data is widely used by AI developers
robots.txt control
Yes
Impact if blocked
Future content excluded from Common Crawl's open archive — previously archived data isn't retroactively removed
How to identify
User-Agent: CCBot

What it does

CCBot's primary purpose is to build Common Crawl's open web archive. Common Crawl publishes this dataset openly, and it is widely used by AI and ML developers as raw training material — rather than being a single operator's proprietary crawl.

Common Crawl states that CCBot respects the robots.txt standard. Because the archive is cumulative, blocking CCBot going forward does not remove content already captured in past archive snapshots.

Why this is optional to block

  • Only affects whether future content enters a shared open dataset that AI developers may draw on
  • Common Crawl documents CCBot as honoring the robots.txt standard
  • Blocking it does not undo any past inclusion in earlier Common Crawl archive snapshots

Identification & robots.txt

User-Agent identifier
CCBot
Verification

Additional verification information is available in Common Crawl's official documentation.

Robots.txt guidance — to opt out
User-agent: CCBot Disallow: /