Skip to content
Training Crawler

Meta-ExternalAgent

Operated by Meta · Trains Meta AI & Llama models

GEO Recommendation: Optional to Block

Meta-ExternalAgent crawls web content for Meta AI model training and product improvement. Meta's documentation states it honors robots.txt, and it is distinct from Meta's user-triggered fetcher, Meta-ExternalFetcher.

Owner
Meta
Trigger
Automated crawl
Used for AI Training
Yes
robots.txt control
Yes
Impact if blocked
Blocks this crawler's access to the disallowed content; user-triggered Meta-ExternalFetcher is controlled separately
How to identify
User-Agent: Meta-ExternalAgent

What it does

Meta-ExternalAgent crawls publicly accessible web content for Meta AI model training and product improvement.

It is a separate, autonomous crawler from Meta-ExternalFetcher, which instead fetches individual pages in response to a specific user request and is not a training crawler.

Why this is optional to block

  • Only affects whether your content contributes to Meta's AI training datasets
  • Has no effect on Meta-ExternalFetcher's user-triggered fetches
  • Meta's documentation states it honors robots.txt, so a Disallow rule is the standard opt-out mechanism

Identification & robots.txt

User-Agent identifier
Meta-ExternalAgent
Verification

Additional verification information is available in Meta's official documentation.

Robots.txt guidance — to opt out of training
User-agent: Meta-ExternalAgent Disallow: /