Skip to content
AI Crawlers & Bots Directory

Which AI bots should access your website?

This directory covers identifiable AI crawlers and user-triggered agents, robots.txt control tokens, and major AI platforms with web access but without a publicly documented bot identity — what each does, and our GEO recommendation where one applies.

24
Bots & tokens tracked
14
Recommended: keep open
10
Optional to block
ChatGPT-User
OpenAI · Powers ChatGPT
AI Assistant

Fetches a page in direct response to a specific ChatGPT user action or request.

Keep Open Profile
OAI-SearchBot
OpenAI · Powers ChatGPT Search
Search Crawler

Crawls content for ChatGPT Search — OpenAI directly ties its access to being found, shown, and cited.

Keep Open Profile
GPTBot
OpenAI · Powers model training
Training Crawler

Collects public content that may be used to train future OpenAI models.

Optional to Block Profile
OAI-AdsBot
OpenAI · Powers ChatGPT Ads
Ads Verification Crawler

Visits landing pages submitted as ChatGPT ads to check policy compliance and relevance — not related to search visibility or model training.

Optional to Block Profile
Claude-User
Anthropic · Powers Claude
AI Assistant

Fetches pages when a Claude user request needs current web content.

Keep Open Profile
Claude-SearchBot
Anthropic · Powers Claude search & citations
Search Crawler

Indexes the web to improve the quality and relevance of Claude's search results.

Keep Open Profile
ClaudeBot
Anthropic · Powers Claude model training
Training Crawler

Collects web content potentially used to train future Claude models.

Optional to Block Profile
Perplexity-User
Perplexity · Powers Perplexity
AI Assistant

May fetch a page directly for a specific user question and cite it in the answer.

Keep Open Profile
PerplexityBot
Perplexity · Powers Perplexity search
Search Crawler

Builds Perplexity's search index; not officially used to train foundation models.

Keep Open Profile
Googlebot
Google · Powers Google Search & AI Overviews
Search Crawler

Google's core search index, which also powers Google's AI features — do not block on a commercial site.

Keep Open Profile
Google-Extended
Google · Governs Gemini training & grounding use
Control Token

A robots control token (not a crawler) governing whether already-crawled content trains and grounds Gemini.

Optional to Block Profile
DuckAssistBot
DuckDuckGo · Powers DuckAssist
Search Crawler

Fetches pages in real time for AI-assisted answers with sources; not used to train models.

Keep Open Profile
Meta-ExternalFetcher
Meta · Supports user-initiated web retrieval in Meta products
AI Assistant

Fetches specific web content on demand to power Meta AI features.

Keep Open Profile
Meta-ExternalAgent
Meta · Trains Meta AI & Llama models
Training Crawler

Collects and indexes content for Meta's AI systems.

Optional to Block Profile
MistralAI-User
Mistral AI · User-triggered retrieval
AI Assistant

User-triggered fetch of web pages for up-to-date AI answers.

Keep Open Profile
MistralAI-Index
Mistral AI · Powers Mistral search functionality
Search Crawler

Automated crawler indexing web content for Mistral AI search functionality.

Keep Open Profile
MistralAI-Training
Mistral AI · Collects data for model training
Training Crawler

Automated crawler collecting web content for Mistral's generative AI training datasets.

Optional to Block Profile
Applebot
Apple · Powers Siri, Spotlight & Safari
Search Crawler

Powers Siri, Spotlight and Safari; now also used as live context for Apple's AI answers.

Keep Open Profile
Applebot-Extended
Apple · Governs Apple Intelligence training use
Control Token

A robots control token governing whether Applebot data trains foundation models.

Optional to Block Profile
Amazonbot
Amazon · Powers Amazon products & AI models
Training Crawler

Collects web content for Amazon services, including Alexa's answer systems.

Optional to Block Profile
Amzn-SearchBot
Amazon · Powers Amazon search experiences, including Alexa
Search Crawler

Used for search experiences like Alexa; Amazon states it is not used to train generative AI models.

Keep Open Profile
Amzn-User
Amazon · Powers user-directed Alexa requests
AI Assistant

Fetches a page on behalf of a user-triggered Alexa/Amazon request for current information.

Keep Open Profile
CCBot
Common Crawl · Feeds a shared, multi-company dataset
Dataset Crawler

Builds the Common Crawl dataset widely used across the ML/LLM ecosystem.

Optional to Block Profile
Bytespider
ByteDance · Widely attributed to ByteDance · Purpose not officially documented
Training Crawler

ByteDance's crawler, associated with data collection for its AI products.

Optional to Block Profile
No Public Bot Identity

AI platforms without a public bot identity

These platforms power widely-used AI assistants, but haven't published a dedicated crawler user-agent or IP range — so their traffic can't be verified or selectively controlled via robots.txt today.

Qwen
Alibaba
AI Platform

Powers Qwen Chat and Quark. Alibaba hasn't published a dedicated bot user-agent or IP range for its web-fetching traffic.

No Bot Identity Profile
DeepSeek
DeepSeek
AI Platform

Powers DeepSeek's chat and search-augmented answers. No public crawler or fetch-bot identity has been documented.

No Bot Identity Profile
Grok
xAI
AI Platform

Answers questions inside X and the Grok app, sometimes citing live web content, without a published bot identity.

No Bot Identity Profile