Home › GEO & AEO › AI crawlers
robots.txt for AI crawlers: GPTBot, ClaudeBot & co.
A full block for search and retrieval crawlers can prevent current content from being fetched as an answer source. Training, search, and user-triggered bots serve different purposes and must be assessed separately.
At a glance
OAI-SearchBot, Claude-SearchBot, and PerplexityBot support search or retrieval. GPTBot, ClaudeBot, and CCBot concern training or datasets. klarava treats only unintended full blocks of controllable search/retrieval bots as a current readiness issue.
What are AI crawlers?
They are bots that AI providers use to read the web, either to train models or to fetch live content for answers. Each identifies itself with a user-agent name you can target in robots.txt.
- GPTBot (OpenAI): training data.
- OAI-SearchBot (OpenAI): content for ChatGPT search.
- Google-Extended (Google): controls use for Gemini/AI, separate from normal Google indexing.
- Claude-SearchBot / Claude-User (Anthropic): search and user-triggered retrieval.
- ClaudeBot (Anthropic): training.
- PerplexityBot (Perplexity).
- CCBot (Common Crawl, used by many AI datasets).
How robots.txt controls them
robots.txt sits at the root of your domain and tells crawlers what they may fetch. You can set rules per user-agent. Reputable AI crawlers honor it.
Example: allow AI crawlers, keep private areas out
User-agent: * Allow: / Disallow: /admin/ Sitemap: https://yourdomain.com/sitemap.xml
A wildcard User-agent: * that allows crawling generally admits supported bots. Remove a specific block only after checking its purpose: a deliberate training block may stay, while an unintended search block should be lifted.
The trade-off
Training and current retrieval are separate decisions. You can block training bots for content or rights reasons while allowing search/retrieval bots. Allowing a bot still does not guarantee a citation.
How klarava helps
klarava's free check classifies full blocks by bot purpose. Only controllable search/retrieval blocks affect the GEO score as a visibility prerequisite; training and dataset policy remain separate.
Check whether AI crawlers can reach your site, for free.
Run the free checkFrequently asked questions
Should I allow AI crawlers?
Decide by purpose. Bots you want for current search and retrieval should not be unintentionally fully blocked. Training and dataset access can be managed independently according to your content and rights policy.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot is OpenAI's crawler used for training data; OAI-SearchBot fetches content to show in ChatGPT search results. You can allow or block them independently in robots.txt.
Does robots.txt guarantee blocking?
robots.txt is a directive that reputable crawlers honor, not a hard technical block. To truly prevent access you need server-side measures. For AI visibility, the relevant point is simply: do not accidentally block the crawlers you want. More on GEO & AEO