GPTBot
Operated by OpenAI · Powers model training
GPTBot collects publicly available web content that may be used to train future OpenAI models. It's a broad, automated crawl unrelated to ChatGPT Search visibility or any specific user's request — blocking it opts your site out of training data collection without affecting search or assistant visibility.
What it does
GPTBot crawls publicly available web pages to collect content that OpenAI says may be used in training future OpenAI models. It operates independently of OAI-SearchBot and ChatGPT-User — blocking GPTBot has no effect on whether your site appears in ChatGPT Search or how ChatGPT-User serves live user requests.
It is the training-data counterpart to OAI-SearchBot's search-index crawl, following the same design pattern as other operators' training bots such as ClaudeBot and Meta-ExternalAgent.
Why this is optional to block
- Only affects whether your content is used to train future OpenAI models
- Has no effect on ChatGPT Search visibility or ChatGPT-User's live fetches
- Many publishers block GPTBot specifically to opt out of training while keeping search/assistant access open
Identification & robots.txt
GPTBot
Additional verification information is available in OpenAI's official documentation.
User-agent: GPTBot
Disallow: /