All GEO insights

AI crawler access: separate search discovery, model training and user-requested visits

AI access is not one setting. Provider-specific controls can distinguish automatic search crawling, content that may be used for future model training, and visits made in response to a user. Google-Extended also covers specified Gemini grounding uses. Robots directives are not authentication, and neither a crawl nor a policy setting guarantees search inclusion, citations or removal from trained models.

In this guide

Can a site owner control AI crawler access separately for search, training and user requests?

Often, a provider documents separate controls for automatic search crawling, content that may be used for model training or grounding, and user-triggered retrieval. Their effects are provider-specific: for example, OpenAI documents OAI-SearchBot, GPTBot and ChatGPT-User separately, while Google says Google-Extended controls specified future Gemini training and grounding uses without affecting Google Search inclusion or ranking. robots.txt provides crawler directives, not access authorization. Check the relevant provider rules and actual page delivery separately; none of these controls guarantees inclusion, a citation, deletion from a trained model or universal compliance.

One page visit does not answer every access question

Providers may use different crawlers or product tokens for automatic search discovery, content that may be used for model development, and retrieval initiated by a person. These are separate purposes, so a choice about one should not be treated as a universal AI-access switch. Google-Extended is a notable distinction: Google documents it as a control for future Gemini model training and specified grounding uses, not as a control for Google Search inclusion or ranking.

Keep access observations distinct from answer outcomes. A crawl, a retrieval, a mention, a citation, a recommendation and a referral are different observations. Evidence of one does not establish another, and the illustration is not a guaranteed sequence.

Six separate panels labelled Crawl, Retrieval, Mention, Citation, Recommendation and Referral illustrate different AI-visibility observations.
These are separate observations: evidence of a crawl, mention or citation does not establish a recommendation or referral. Conceptual illustration.View full-size illustration

Sources: Overview of OpenAI Crawlers; Does Anthropic crawl data from the web, and how can site owners block the crawler?; List of Google common crawlers; AI features and your website

Provider controls and their documented boundaries

The matrix summarizes the stated purpose of each control, not a universal rule for every bot or site. A user-agent string in an HTTP request is not always the same thing as the product token used in a robots.txt group. Google-Extended has a robots.txt control token but no separate HTTP request user-agent. Check provider documentation for the host and use in question; the documented choices and exceptions are not interchangeable.

Documented controls and important qualifications

Scroll to read all columns.

Provider controlDocumented purposeImportant boundary
OpenAI OAI-SearchBotAutomatic discovery for ChatGPT search features.OpenAI says sites opted out of OAI-SearchBot will not appear in ChatGPT search answers, though they may still appear as navigational links.
OpenAI GPTBotCrawls content that may be used to train OpenAI foundation models.OpenAI describes GPTBot as independent of OAI-SearchBot. Disallowing GPTBot signals that content should not be used for this training purpose; it does not establish removal from already trained models.
OpenAI ChatGPT-UserCertain visits triggered by a ChatGPT or Custom GPT user; not automatic web crawling.OpenAI says robots.txt rules may not apply to these user-initiated requests. This control does not manage Search opt-outs; OAI-SearchBot is the documented Search control.
Anthropic Claude-SearchBotSearch-result quality and indexing for user search.Anthropic says disabling it prevents indexing for search optimization and may reduce visibility or accuracy in user search results.
Anthropic ClaudeBotCollects web content that could potentially contribute to model training.Anthropic says restricting access signals that future materials should be excluded from its AI model training datasets; it is not a claim about previously trained models.
Anthropic Claude-UserRetrieves web content in response to a user query.Anthropic says disabling it prevents this retrieval and may reduce visibility for user-directed web search. Anthropic states that its bots honor robots.txt directives.
Google GooglebotGoogle Search crawling, including AI Overviews and AI Mode.Google identifies Googlebot as the Search access control. Treating it as training-only would also affect Search crawling; crawling permission alone does not guarantee a search appearance.
Google Google-ExtendedControls whether crawled content may be used for future Gemini model training and specified grounding uses in Gemini Apps and Grounding with Google Search on Vertex AI.It does not affect Google Search inclusion or ranking, and has no separate HTTP request user-agent. Google documents the token as a control for future use, not as removal from models already trained.

Sources: Overview of OpenAI Crawlers; Does Anthropic crawl data from the web, and how can site owners block the crawler?; List of Google common crawlers; AI features and your website

robots.txt is a crawler directive, not a lock

The Robots Exclusion Protocol asks automatic crawlers to honor rules about accessing paths. RFC 9309 explicitly says those rules are not access authorization. A robots.txt file is publicly readable; listing a sensitive path can make it easier to discover, not private. Use application-level access controls, such as authentication, when content must be restricted.

Keep policy text, authentication and actual delivery as separate checks. A page can be blocked from a crawler by hosting or CDN behavior even when the robots rule permits access, and a robots directive does not itself prevent a request. Provider documentation describes specific compliance behavior, but a directive should not be treated as a guarantee that every client will comply. A bot name in a request is not, on its own, proof of identity.

Sources: RFC 9309: Robots Exclusion Protocol; List of Google common crawlers; Overview of OpenAI Crawlers; Does Anthropic crawl data from the web, and how can site owners block the crawler?

A review workflow before changing a setting

Use this sequence to clarify what a proposed change is meant to do. It does not change a site setting or determine whether a page will appear in search or an answer.

  1. Name the content and purpose

    Identify the specific public pages and whether the goal concerns search discovery, possible training use, grounding or user-requested retrieval. Do not treat “AI access” as a single purpose.

  2. Check the matching provider rule

    Read the provider’s documentation and the robots.txt groups for the relevant host. Review existing rules rather than copying a blanket block; check whether the documented control is an HTTP user-agent or a robots product token.

  3. Check actual delivery and protection separately

    Check whether the relevant page and robots.txt can actually be retrieved, including any hosting or CDN restriction. Keep private or account-only material behind proper access controls; a robots rule is not a substitute.

  4. Record what you observed, not what you assume

    Record the rule reviewed and any observed access result. A fetch does not establish search inclusion or a citation. A training opt-out signal does not establish deletion from a model already trained, and one provider’s documented behavior does not establish universal compliance.

Sources: RFC 9309: Robots Exclusion Protocol; AI features and your website; Overview of OpenAI Crawlers; Does Anthropic crawl data from the web, and how can site owners block the crawler?

Sources

Read the JSON companion

Read next

Review controls by purpose, not by the word “AI”

Name the public content and the specific search, training or grounding purpose first. Then check the provider’s current documentation and the applicable host rules. Keep robots directives separate from authentication and actual page delivery. Treat a recorded fetch as evidence of a visit, not proof of inclusion or a citation.

Check my brand