Promptwatch Logo

Google-Extended

Google-Extended is a robots. txt product token, not a separate crawler that identifies itself in HTTP requests.
GoogleGoogle-Extended
AI CrawlerAI Training

What is Google-Extended?

Google-Extended is a robots.txt product token, not a separate crawler that identifies itself in HTTP requests. Google fetches pages with existing crawler user agents, then uses the Google-Extended group as a control over whether the crawled content may be used by specified generative AI products.

Google says the token covers training future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini. It also controls grounding with content from the Google Search index in Gemini Apps and Grounding with Google Search on Vertex AI. A disallowed page is outside those Google-Extended uses.

This setting is independent of ordinary Google Search. Google states that a Google-Extended rule does not affect inclusion in Search and is not a Search ranking signal. Publishers can therefore keep pages available to Googlebot for search indexing while declining the uses covered by Google-Extended.

There is no Google-Extended user-agent string to find in an access log. Google's common crawler documentation says requests are made with existing Google user agents. A request from Googlebot cannot, by itself, reveal which later Google-Extended use might apply to the fetched content.

Google's common crawlers automatically obey robots.txt, and the directory records Google-Extended as respecting it. The control can cover the entire site or selected paths. It is a use preference for content Google crawls, not a firewall rule that blocks every Google request to those URLs.

Because Google-Extended is not an HTTP identity, a dedicated IP check or signature for that token would not make sense. Site owners who need to verify an underlying Google crawler should use Google's published crawler-verification methods, but should keep that identity question separate from the Google-Extended content-use policy.

Relevant for AI searchGeminiAI Overviews

Is Google-Extended relevant for AI search?

Yes. Google-Extended feeds Gemini, AI Overviews, so the pages it can reach shape what those AI products say about you.

Google-Extended gathers public web content that can end up in the training data for large language models. Once your pages are in that set, they influence how the operator's models talk about you for that model generation. Allowing it lets your own writing carry weight in those answers; blocking it means the models learn about you from third parties instead.

Track Google-Extended with Promptwatch

Promptwatch classifies Google-Extended (Google-Extended) in real time from your server and CDN logs. See exactly when it visits, which pages it requests, the status codes it gets, and how those crawls map to AI citations.

How to handle Google-Extended

If you do not want crawled content used for the Gemini training and grounding purposes covered by Google-Extended, publish this group:

User-agent: Google-Extended
Disallow: /

Use path-specific Disallow values if only part of the site needs the restriction. Do not block Googlebot merely to set this preference, because Google says the Google-Extended group has no effect on Search inclusion or ranking.

Do not build a firewall rule that expects an HTTP user agent named Google-Extended. Google documents it only as a robots.txt product token; the fetch itself uses an existing Google crawler identity.

Examples

  • A news site disallows `Google-Extended` across its articles while leaving `Googlebot` access unchanged for ordinary Search indexing.
  • A developer portal allows Google-Extended on public API references but excludes a path containing licensed technical manuals.
  • An analyst looking at Googlebot requests does not label any one request as Google-Extended traffic, because the product token never appears as a separate HTTP user agent.

Frequently asked questions about Google-Extended

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

No. Google defines it as a standalone product token used in robots.txt. Existing Google crawler user agents perform the underlying fetches.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard