Google-Extended
Operated by Google · Governs Gemini training & grounding use
Google-Extended is not a crawler and has no bot of its own — it's a standalone robots.txt token. Content is fetched by Googlebot as usual; Google-Extended only controls whether that already-crawled content may be used to train Gemini and Vertex AI generative models.
What it does
Google-Extended doesn't have its own HTTP user agent or crawler — Google's documentation is explicit that crawling is done with existing Google user-agent strings, and the robots.txt token is used only in a control capacity over how that data is later used.
Disallowing Google-Extended governs whether content already crawled by Googlebot may be used for training future Gemini models, for grounding in Gemini Apps, and for Grounding with Google Search on Vertex AI. Google states this token does not impact a site's inclusion or ranking in Google Search.
Why this is optional to block
- It's a training-use switch, not an access control — disallowing it doesn't stop Googlebot from crawling or indexing your site
- Google states blocking it has no effect on Search inclusion or ranking
- Whether to allow your content into Gemini/Vertex AI training is a content-strategy decision, not a technical requirement
Identification & robots.txt
Google-Extended
Google-Extended has no separate HTTP user-agent to verify — it's a named group in your robots.txt file. Additional detail is available in Google's official documentation.
User-agent: Google-Extended
Disallow: /