Which AI bots should access your website?
This directory covers identifiable AI crawlers and user-triggered agents, robots.txt control tokens, and major AI platforms with web access but without a publicly documented bot identity — what each does, and our GEO recommendation where one applies.
No entries match the current search and filters.
Fetches a page in direct response to a specific ChatGPT user action or request.
Crawls content for ChatGPT Search — OpenAI directly ties its access to being found, shown, and cited.
Collects public content that may be used to train future OpenAI models.
Visits landing pages submitted as ChatGPT ads to check policy compliance and relevance — not related to search visibility or model training.
Fetches pages when a Claude user request needs current web content.
Indexes the web to improve the quality and relevance of Claude's search results.
Collects web content potentially used to train future Claude models.
May fetch a page directly for a specific user question and cite it in the answer.
Builds Perplexity's search index; not officially used to train foundation models.
Google's core search index, which also powers Google's AI features — do not block on a commercial site.
A robots control token (not a crawler) governing whether already-crawled content trains and grounds Gemini.
Fetches pages in real time for AI-assisted answers with sources; not used to train models.
Fetches specific web content on demand to power Meta AI features.
Collects and indexes content for Meta's AI systems.
User-triggered fetch of web pages for up-to-date AI answers.
Automated crawler indexing web content for Mistral AI search functionality.
Automated crawler collecting web content for Mistral's generative AI training datasets.
Powers Siri, Spotlight and Safari; now also used as live context for Apple's AI answers.
A robots control token governing whether Applebot data trains foundation models.
Collects web content for Amazon services, including Alexa's answer systems.
Used for search experiences like Alexa; Amazon states it is not used to train generative AI models.
Fetches a page on behalf of a user-triggered Alexa/Amazon request for current information.
Builds the Common Crawl dataset widely used across the ML/LLM ecosystem.
ByteDance's crawler, associated with data collection for its AI products.
AI platforms without a public bot identity
These platforms power widely-used AI assistants, but haven't published a dedicated crawler user-agent or IP range — so their traffic can't be verified or selectively controlled via robots.txt today.
Powers Qwen Chat and Quark. Alibaba hasn't published a dedicated bot user-agent or IP range for its web-fetching traffic.
Powers DeepSeek's chat and search-augmented answers. No public crawler or fetch-bot identity has been documented.
Answers questions inside X and the Grok app, sometimes citing live web content, without a published bot identity.