{"schemaVersion":"ready-scan.discovery-resource.v1","id":"urn:readyscan:discovery:ai-crawler-search-training-controls","slug":"ai-crawler-search-training-controls","kind":"guide","inLanguage":"en","canonicalUrl":"https://readyscan.ai/insights/ai-crawler-search-training-controls","machineUrl":"https://readyscan.ai/discovery/resources/ai-crawler-search-training-controls.json","title":"AI crawler controls: search, training and user-requested visits | Ready Scan","description":"A practical comparison of provider-specific crawler controls, what robots.txt can and cannot do, and a review workflow that does not promise indexing or model unlearning.","h1":"AI crawler access: separate search discovery, model training and user-requested visits","summary":"AI access is not one setting. Provider-specific controls can distinguish automatic search crawling, content that may be used for future model training, and visits made in response to a user. Google-Extended also covers specified Gemini grounding uses. Robots directives are not authentication, and neither a crawl nor a policy setting guarantees search inclusion, citations or removal from trained models.","answer":{"question":"Can a site owner control AI crawler access separately for search, training and user requests?","text":"Often, a provider documents separate controls for automatic search crawling, content that may be used for model training or grounding, and user-triggered retrieval. Their effects are provider-specific: for example, OpenAI documents OAI-SearchBot, GPTBot and ChatGPT-User separately, while Google says Google-Extended controls specified future Gemini training and grounding uses without affecting Google Search inclusion or ranking. robots.txt provides crawler directives, not access authorization. Check the relevant provider rules and actual page delivery separately; none of these controls guarantees inclusion, a citation, deletion from a trained model or universal compliance.","sourceIds":["openai-bots","anthropic-bots","google-crawlers","google-ai-features","robots-standard"]},"sections":[{"title":"One page visit does not answer every access question","body":"Providers may use different crawlers or product tokens for automatic search discovery, content that may be used for model development, and retrieval initiated by a person. These are separate purposes, so a choice about one should not be treated as a universal AI-access switch. Google-Extended is a notable distinction: Google documents it as a control for future Gemini model training and specified grounding uses, not as a control for Google Search inclusion or ranking.\n\nKeep access observations distinct from answer outcomes. A crawl, a retrieval, a mention, a citation, a recommendation and a referral are different observations. Evidence of one does not establish another, and the illustration is not a guaranteed sequence.","sourceIds":["openai-bots","anthropic-bots","google-crawlers","google-ai-features"],"figure":{"src":"/assets/insights/ai-visibility-signals-1280.webp","width":1280,"height":720,"alt":"Six separate panels labelled Crawl, Retrieval, Mention, Citation, Recommendation and Referral illustrate different AI-visibility observations.","caption":"These are separate observations: evidence of a crawl, mention or citation does not establish a recommendation or referral. Conceptual illustration.","variants":[{"src":"/assets/insights/ai-visibility-signals-640.webp","width":640}]}},{"title":"Provider controls and their documented boundaries","body":"The matrix summarizes the stated purpose of each control, not a universal rule for every bot or site. A user-agent string in an HTTP request is not always the same thing as the product token used in a robots.txt group. Google-Extended has a robots.txt control token but no separate HTTP request user-agent. Check provider documentation for the host and use in question; the documented choices and exceptions are not interchangeable.","sourceIds":["openai-bots","anthropic-bots","google-crawlers","google-ai-features"],"table":{"caption":"Documented controls and important qualifications","columns":["Provider control","Documented purpose","Important boundary"],"rows":[["OpenAI OAI-SearchBot","Automatic discovery for ChatGPT search features.","OpenAI says sites opted out of OAI-SearchBot will not appear in ChatGPT search answers, though they may still appear as navigational links."],["OpenAI GPTBot","Crawls content that may be used to train OpenAI foundation models.","OpenAI describes GPTBot as independent of OAI-SearchBot. Disallowing GPTBot signals that content should not be used for this training purpose; it does not establish removal from already trained models."],["OpenAI ChatGPT-User","Certain visits triggered by a ChatGPT or Custom GPT user; not automatic web crawling.","OpenAI says robots.txt rules may not apply to these user-initiated requests. This control does not manage Search opt-outs; OAI-SearchBot is the documented Search control."],["Anthropic Claude-SearchBot","Search-result quality and indexing for user search.","Anthropic says disabling it prevents indexing for search optimization and may reduce visibility or accuracy in user search results."],["Anthropic ClaudeBot","Collects web content that could potentially contribute to model training.","Anthropic says restricting access signals that future materials should be excluded from its AI model training datasets; it is not a claim about previously trained models."],["Anthropic Claude-User","Retrieves web content in response to a user query.","Anthropic says disabling it prevents this retrieval and may reduce visibility for user-directed web search. Anthropic states that its bots honor robots.txt directives."],["Google Googlebot","Google Search crawling, including AI Overviews and AI Mode.","Google identifies Googlebot as the Search access control. Treating it as training-only would also affect Search crawling; crawling permission alone does not guarantee a search appearance."],["Google Google-Extended","Controls whether crawled content may be used for future Gemini model training and specified grounding uses in Gemini Apps and Grounding with Google Search on Vertex AI.","It does not affect Google Search inclusion or ranking, and has no separate HTTP request user-agent. Google documents the token as a control for future use, not as removal from models already trained."]]}},{"title":"robots.txt is a crawler directive, not a lock","body":"The Robots Exclusion Protocol asks automatic crawlers to honor rules about accessing paths. RFC 9309 explicitly says those rules are not access authorization. A robots.txt file is publicly readable; listing a sensitive path can make it easier to discover, not private. Use application-level access controls, such as authentication, when content must be restricted.\n\nKeep policy text, authentication and actual delivery as separate checks. A page can be blocked from a crawler by hosting or CDN behavior even when the robots rule permits access, and a robots directive does not itself prevent a request. Provider documentation describes specific compliance behavior, but a directive should not be treated as a guarantee that every client will comply. A bot name in a request is not, on its own, proof of identity.","sourceIds":["robots-standard","google-crawlers","openai-bots","anthropic-bots"]},{"title":"A review workflow before changing a setting","body":"Use this sequence to clarify what a proposed change is meant to do. It does not change a site setting or determine whether a page will appear in search or an answer.","sourceIds":["robots-standard","google-ai-features","openai-bots","anthropic-bots"],"items":[{"title":"Name the content and purpose","body":"Identify the specific public pages and whether the goal concerns search discovery, possible training use, grounding or user-requested retrieval. Do not treat “AI access” as a single purpose."},{"title":"Check the matching provider rule","body":"Read the provider’s documentation and the robots.txt groups for the relevant host. Review existing rules rather than copying a blanket block; check whether the documented control is an HTTP user-agent or a robots product token."},{"title":"Check actual delivery and protection separately","body":"Check whether the relevant page and robots.txt can actually be retrieved, including any hosting or CDN restriction. Keep private or account-only material behind proper access controls; a robots rule is not a substitute."},{"title":"Record what you observed, not what you assume","body":"Record the rule reviewed and any observed access result. A fetch does not establish search inclusion or a citation. A training opt-out signal does not establish deletion from a model already trained, and one provider’s documented behavior does not establish universal compliance."}]}],"sources":[{"id":"openai-bots","title":"Overview of OpenAI Crawlers","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/bots","accessedOn":"2026-10-10"},{"id":"anthropic-bots","title":"Does Anthropic crawl data from the web, and how can site owners block the crawler?","publisher":"Anthropic","url":"https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler","accessedOn":"2026-10-10"},{"id":"google-crawlers","title":"List of Google common crawlers","publisher":"Google for Developers","url":"https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers","accessedOn":"2026-10-10"},{"id":"google-ai-features","title":"AI features and your website","publisher":"Google Search Central","url":"https://developers.google.com/search/docs/appearance/ai-features","accessedOn":"2026-10-06"},{"id":"robots-standard","title":"RFC 9309: Robots Exclusion Protocol","publisher":"IETF","url":"https://www.rfc-editor.org/rfc/rfc9309.html","accessedOn":"2026-10-10"}],"ownership":{"state":"ready_owned","publisher":{"name":"Ready Scan","url":"https://readyscan.ai"},"verificationMethod":"first_party_publication","verificationSource":"https://readyscan.ai/insights/ai-crawler-search-training-controls"},"datePublished":"2026-10-10","dateModified":"2026-10-10","related":[{"id":"urn:readyscan:discovery:ai-citation-visibility","canonicalUrl":"https://readyscan.ai/ai-citation-visibility","machineUrl":"https://readyscan.ai/discovery/resources/ai-citation-visibility.json"},{"id":"urn:readyscan:discovery:website-updates-search-ai-answers","canonicalUrl":"https://readyscan.ai/insights/website-updates-search-ai-answers","machineUrl":"https://readyscan.ai/discovery/resources/website-updates-search-ai-answers.json"},{"id":"urn:readyscan:discovery:geo-aeo-seo","canonicalUrl":"https://readyscan.ai/insights/geo-aeo-seo","machineUrl":"https://readyscan.ai/discovery/resources/geo-aeo-seo.json"},{"id":"urn:readyscan:source:source-2d8d883ba8f289ac","canonicalUrl":"https://readyscan.ai/sources/source-2d8d883ba8f289ac","machineUrl":"https://readyscan.ai/discovery/sources/source-2d8d883ba8f289ac.json"},{"id":"urn:readyscan:source:source-0e8eaf1d95acd8d8","canonicalUrl":"https://readyscan.ai/sources/source-0e8eaf1d95acd8d8","machineUrl":"https://readyscan.ai/discovery/sources/source-0e8eaf1d95acd8d8.json"},{"id":"urn:readyscan:source:source-8f6bb09bbba1b2d8","canonicalUrl":"https://readyscan.ai/sources/source-8f6bb09bbba1b2d8","machineUrl":"https://readyscan.ai/discovery/sources/source-8f6bb09bbba1b2d8.json"},{"id":"urn:readyscan:source:source-0a5f4e0044c75068","canonicalUrl":"https://readyscan.ai/sources/source-0a5f4e0044c75068","machineUrl":"https://readyscan.ai/discovery/sources/source-0a5f4e0044c75068.json"},{"id":"urn:readyscan:source:source-d871053673d1c958","canonicalUrl":"https://readyscan.ai/sources/source-d871053673d1c958","machineUrl":"https://readyscan.ai/discovery/sources/source-d871053673d1c958.json"}],"structuredData":{"@context":"https://schema.org","@type":"Article","headline":"AI crawler controls: search, training and user-requested visits | Ready Scan","description":"A practical comparison of provider-specific crawler controls, what robots.txt can and cannot do, and a review workflow that does not promise indexing or model unlearning.","mainEntityOfPage":"https://readyscan.ai/insights/ai-crawler-search-training-controls","datePublished":"2026-10-10","dateModified":"2026-10-10","publisher":{"@type":"Organization","name":"Ready Scan","url":"https://readyscan.ai"},"author":{"@type":"Organization","name":"Ready Scan","url":"https://readyscan.ai"},"citation":["https://developers.openai.com/api/docs/bots","https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler","https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers","https://developers.google.com/search/docs/appearance/ai-features","https://www.rfc-editor.org/rfc/rfc9309.html"],"image":{"@type":"ImageObject","url":"https://readyscan.ai/assets/insights/ai-visibility-signals-1280.webp","width":1280,"height":720,"caption":"These are separate observations: evidence of a crawl, mention or citation does not establish a recommendation or referral. Conceptual illustration.","description":"Six separate panels labelled Crawl, Retrieval, Mention, Citation, Recommendation and Referral illustrate different AI-visibility observations."},"inLanguage":"en"}}