Skip to content
Training Crawler

GPTBot

Operated by OpenAI · Powers model training

GEO Recommendation: Optional to Block

GPTBot collects publicly available web content that may be used to train future OpenAI models. It's a broad, automated crawl unrelated to ChatGPT Search visibility or any specific user's request — blocking it opts your site out of training data collection without affecting search or assistant visibility.

Owner
OpenAI
Trigger
Automated (continuous crawl)
Used for AI Training
Yes
robots.txt control
Yes
Impact if blocked
Site excluded from future OpenAI model training data — no effect on ChatGPT Search or assistant visibility
How to identify
User-Agent: GPTBot

What it does

GPTBot crawls publicly available web pages to collect content that OpenAI says may be used in training future OpenAI models. It operates independently of OAI-SearchBot and ChatGPT-User — blocking GPTBot has no effect on whether your site appears in ChatGPT Search or how ChatGPT-User serves live user requests.

It is the training-data counterpart to OAI-SearchBot's search-index crawl, following the same design pattern as other operators' training bots such as ClaudeBot and Meta-ExternalAgent.

Why this is optional to block

  • Only affects whether your content is used to train future OpenAI models
  • Has no effect on ChatGPT Search visibility or ChatGPT-User's live fetches
  • Many publishers block GPTBot specifically to opt out of training while keeping search/assistant access open

Identification & robots.txt

User-Agent identifier
GPTBot
Verification

Additional verification information is available in OpenAI's official documentation.

Robots.txt guidance — to opt out of training
User-agent: GPTBot Disallow: /