The Four Kinds of AI Bots
Not every AI crawler wants the same thing from your page. The generator sorts its catalogue of over 50 AI user agents into four roles, because blocking each role has a different cost.
Training data collectors fetch pages to build datasets for future models. GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended, CCBot (Common Crawl) and Bytespider (ByteDance) belong here, along with newer names such as DeepSeekBot and MistralAI-Training. Blocking them keeps your text out of future training runs and has no effect on search visibility.
AI search indexers build the indexes that AI answer engines draw on and cite. OAI-SearchBot, Claude-SearchBot, PerplexityBot and DuckAssistBot are typical. Block these and your pages stop appearing as sources in AI search answers.
User-triggered fetchers visit only when a person asks an assistant to open a specific URL: ChatGPT-User, Claude-User, Perplexity-User and Google-NotebookLM, for example. Blocking them means a reader who pastes your link into a chat gets a refusal instead of a summary.
Browsing agents such as ChatGPT Agent, Operator, GoogleAgent-Mariner and Amazon's NovaAct navigate sites, click and fill forms on someone's behalf.
OpenAI shows most clearly why the split matters, because it runs a separate bot for each job:
| Token | Role | What blocking it costs |
|---|---|---|
GPTBot |
Training | Nothing in search; your pages leave future training data |
OAI-SearchBot |
AI search | Your site stops being cited in ChatGPT search |
ChatGPT-User |
User-triggered | ChatGPT cannot open your page when a user asks |
Blocking GPTBot alone is the common middle ground: you opt out of training and stay citable. That is exactly what the default Training crawlers only preset does across every operator in the list.