klaravaGEO/AEO report
DE EN

Home › GEO & AEO › AI crawlers

robots.txt for AI crawlers: GPTBot, ClaudeBot & co.

A full block for search and retrieval crawlers can prevent current content from being fetched as an answer source. Training, search, and user-triggered bots serve different purposes and must be assessed separately.

At a glance

OAI-SearchBot, Claude-SearchBot, and PerplexityBot support search or retrieval. GPTBot, ClaudeBot, and CCBot concern training or datasets. klarava treats only unintended full blocks of controllable search/retrieval bots as a current readiness issue.

What are AI crawlers?

They are bots that AI providers use to read the web, either to train models or to fetch live content for answers. Each identifies itself with a user-agent name you can target in robots.txt.

How robots.txt controls them

robots.txt sits at the root of your domain and tells crawlers what they may fetch. You can set rules per user-agent. Reputable AI crawlers honor it.

Example: allow AI crawlers, keep private areas out

User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://yourdomain.com/sitemap.xml

A wildcard User-agent: * that allows crawling generally admits supported bots. Remove a specific block only after checking its purpose: a deliberate training block may stay, while an unintended search block should be lifted.

The trade-off

Training and current retrieval are separate decisions. You can block training bots for content or rights reasons while allowing search/retrieval bots. Allowing a bot still does not guarantee a citation.

How klarava helps

klarava's free check classifies full blocks by bot purpose. Only controllable search/retrieval blocks affect the GEO score as a visibility prerequisite; training and dataset policy remain separate.

Check whether AI crawlers can reach your site, for free.

Run the free check

Frequently asked questions

Should I allow AI crawlers?

Decide by purpose. Bots you want for current search and retrieval should not be unintentionally fully blocked. Training and dataset access can be managed independently according to your content and rights policy.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot is OpenAI's crawler used for training data; OAI-SearchBot fetches content to show in ChatGPT search results. You can allow or block them independently in robots.txt.

Does robots.txt guarantee blocking?

robots.txt is a directive that reputable crawlers honor, not a hard technical block. To truly prevent access you need server-side measures. For AI visibility, the relevant point is simply: do not accidentally block the crawlers you want. More on GEO & AEO