CCBot
Operated by Common Crawl · Feeds a shared, multi-company dataset
CCBot is the crawler for Common Crawl, a nonprofit that publishes an open web archive. Unlike single-company bots, this archive is downloaded and used for training by many different AI labs — blocking CCBot is a general opt-out from that shared dataset, not from any one company's models.
What it does
CCBot's primary purpose is to build Common Crawl's open web archive. Common Crawl publishes this dataset openly, and it is widely used by AI and ML developers as raw training material — rather than being a single operator's proprietary crawl.
Common Crawl states that CCBot respects the robots.txt standard. Because the archive is cumulative, blocking CCBot going forward does not remove content already captured in past archive snapshots.
Why this is optional to block
- Only affects whether future content enters a shared open dataset that AI developers may draw on
- Common Crawl documents CCBot as honoring the robots.txt standard
- Blocking it does not undo any past inclusion in earlier Common Crawl archive snapshots
Identification & robots.txt
CCBot
Additional verification information is available in Common Crawl's official documentation.
User-agent: CCBot
Disallow: /