GEO
Google-Extended vs Googlebot: Why They're Not the Same
Googlebot and Google-Extended are two different user-agent tokens Google reads in robots.txt: Googlebot controls whether your site shows up in Google Search, and Google-Extended controls whether your content trains Gemini and feeds AI Overviews. Disallowing one has no effect on the other.
What does Googlebot control?
Googlebot is the crawler behind ordinary Google Search. If Googlebot can fetch a page, that page becomes eligible for indexing and ranking in regular search results. Disallowing Googlebot in robots.txt — or a site-wide Disallow: / under User-agent: * that catches it too — removes a site from Google Search entirely.
What does Google-Extended control instead?
Google-Extended is a separate, narrower token. It doesn't affect Search indexing at all — it only controls whether Google can use a site's content to train Gemini models and to ground AI Overview answers. A site can disallow Google-Extended while Googlebot keeps crawling and ranking it normally.
Why do people mix these up?
Both tokens start with "Google," both live in the same robots.txt file, and both relate to "Google using my content." The mistake usually goes one of two directions:
- Blocking Googlebot thinking it only stops AI training — this actually removes the site from Google Search results entirely, a far bigger consequence than intended.
- Leaving Google-Extended allowed by default, unaware it exists — the site's content keeps feeding Gemini training and AI Overviews even if the owner would have preferred to opt out.
How should you configure both, if you want AI visibility?
For a site trying to maximize GEO visibility — appearing in AI Overviews and being usable as grounding context for Gemini — the answer is to leave both allowed. For a site that wants Search traffic but not to feed AI training, the split configuration looks like this:
| Goal | Googlebot | Google-Extended |
|---|---|---|
| Rank in Google Search + appear in AI Overviews | Allow | Allow |
| Rank in Google Search, opt out of AI training | Allow | Disallow |
| Opt out of Google entirely | Disallow | Irrelevant — already excluded |
Check which of these two tokens your current robots.txt actually sets with the free AI Crawler Checker — most sites have never explicitly set Google-Extended at all, which defaults to allowed.
Frequently Asked Questions
Does blocking Google-Extended remove me from Google Search?
No. Google-Extended only controls whether Google can use your content for Gemini and AI Overviews training/grounding. Googlebot, which controls Search indexing and ranking, is a completely separate user-agent and is unaffected.
Does blocking Googlebot stop me from appearing in AI Overviews?
Yes, indirectly — AI Overviews draw from Google's regular search index, so if Googlebot can't crawl your pages, they can't be indexed or surfaced in an AI Overview either, regardless of your Google-Extended setting.
Can I allow Search but opt out of AI training?
Yes. Allow Googlebot as normal and add a separate Disallow rule for Google-Extended. This keeps your pages indexable and rankable in Google Search while excluding them from the training/grounding data used for Gemini and AI Overviews.
Is Google-Extended the same thing as an AI crawler block?
It's one specific AI-use control among several — GPTBot, ClaudeBot, and PerplexityBot each have their own separate user-agent tokens, so blocking Google-Extended has no effect on those other companies' crawlers.
Related Free Tools
Related Posts
GPTBot vs OAI-SearchBot: What Each One Actually Does
llms.txt in 2026: Does It Actually Do Anything for SEO?