Choose language

AI Crawler Robots.txt Generator

Build a robots.txt that blocks or allows 53 AI crawlers by role, with Content Signals for search, AI input, and AI training.

AI Crawler Robots.txt GeneratorHow it works ↓
Index the page and show links with short excerpts
Use the page as source material for generated answers
Use the page to train or fine-tune AI models
Cloudflare extension: how far a bot may go with content it fetched
The legal wording that gives the signals effect under EU copyright rules
Blocking these has no effect on search visibility
Blocking these removes you from AI search answers and their citations
Blocking these stops assistants from opening your pages on request
Agents that click and fill forms on behalf of a user
Use / for the whole site or a folder such as /articles/
Needed only for the Sitemap and llms.txt pointers
Added as a comment; robots.txt has no llms.txt directive

robots.txt and Content-Signal state preferences; they do not block anything by themselves. Well-behaved crawlers honour them, others ignore them, and Google has said it does not act on Content-Signal. Pair them with server-side bot rules if you need enforcement.

Mehmet Demiray Published Updated
Share

The Four Kinds of AI Bots

Not every AI crawler wants the same thing from your page. The generator sorts its catalogue of over 50 AI user agents into four roles, because blocking each role has a different cost.

Training data collectors fetch pages to build datasets for future models. GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended, CCBot (Common Crawl) and Bytespider (ByteDance) belong here, along with newer names such as DeepSeekBot and MistralAI-Training. Blocking them keeps your text out of future training runs and has no effect on search visibility.

AI search indexers build the indexes that AI answer engines draw on and cite. OAI-SearchBot, Claude-SearchBot, PerplexityBot and DuckAssistBot are typical. Block these and your pages stop appearing as sources in AI search answers.

User-triggered fetchers visit only when a person asks an assistant to open a specific URL: ChatGPT-User, Claude-User, Perplexity-User and Google-NotebookLM, for example. Blocking them means a reader who pastes your link into a chat gets a refusal instead of a summary.

Browsing agents such as ChatGPT Agent, Operator, GoogleAgent-Mariner and Amazon's NovaAct navigate sites, click and fill forms on someone's behalf.

OpenAI shows most clearly why the split matters, because it runs a separate bot for each job:

Token Role What blocking it costs
GPTBot Training Nothing in search; your pages leave future training data
OAI-SearchBot AI search Your site stops being cited in ChatGPT search
ChatGPT-User User-triggered ChatGPT cannot open your page when a user asks

Blocking GPTBot alone is the common middle ground: you opt out of training and stay citable. That is exactly what the default Training crawlers only preset does across every operator in the list.

How the Generated File Works

Press Generate robots.txt and the tool assembles a plain text file in a fixed order. It is built in your browser from the options on screen, nothing is uploaded, and you can copy the result or download it as robots.txt.

  1. Policy comment block (optional). When at least one content signal is set and the policy box is ticked, the file opens with the Content Signals Policy text as # comment lines.
  2. The wildcard group. User-agent: * comes next, then the Content-Signal line if you set any signal, then Allow: /. This group covers every crawler that has no group of its own, so ordinary search engines keep full access.
  3. One group per blocked bot. Each selected crawler gets its own short group, for example User-agent: GPTBot followed by Disallow: /. Groups follow catalogue order and are separated by blank lines.
  4. Pointers (optional). If you enter a site URL, a Sitemap: line and an llms.txt comment can close the file.

The disallow path is shared. Whatever you type in Disallow path (the default / means the whole site) is written under every blocked bot, so one run cannot give GPTBot one path and CCBot another. The path has to start with a slash or the form shows an error.

Why separate groups instead of extra rules inside the wildcard group? Under the Robots Exclusion Protocol a crawler obeys only the most specific group that names it. GPTBot reads its own group and ignores User-agent: *, so the Allow: / there does not cancel its block, and crawlers without a named group are not affected by the blocks.

Above the file, a summary card shows how many AI crawlers are blocked, the Content-Signal values written, which operators are affected (OpenAI, Anthropic, Google and so on) and the number of User-agent groups. That last figure is always the blocked count plus one for the wildcard group, a quick sanity check before you upload.

Content Signals Explained

Content-Signal is a newer robots.txt line introduced through Cloudflare's Content Signals Policy. Instead of telling a bot to stay away, it states what the content may be used for once it has been fetched. It has three keys:

  • search: building a search index and showing links with short excerpts. The policy text says this does not include AI-generated search summaries.
  • ai-input: feeding the page into an AI model at answer time, as in retrieval augmented generation or grounding.
  • ai-train: training or fine-tuning AI models.

Each key can be set to Allow (yes), Disallow (no) or No preference. No preference does not write no; it leaves the key out of the line entirely, which under the policy means you neither grant nor restrict that use. The defaults are search and AI input allowed, AI training disallowed, which produces Content-Signal: search=yes, ai-input=yes, ai-train=no. If every key is left at No preference and no use value is chosen, the line is not written at all.

Permitted use after crawling adds Cloudflare's use parameter to the same line. Its options are immediate (no storage or reuse), reference (index, excerpt and link back) and full (summaries and reproduction allowed). A file with search allowed, training refused and reference use would carry Content-Signal: search=yes, ai-train=no, use=reference.

The optional policy comment block is there because a bare signal is just a token. The comment text spells out what yes, no and a missing value mean, defines each signal, and states that any restrictions are express reservations of rights under Article 4 of EU Directive 2019/790 on copyright in the Digital Single Market. That article allows text and data mining unless the rights holder reserves it in a machine-readable way, and the signal plus the comment are meant to serve as that reservation. How much legal weight it carries in a given case is a question for a lawyer, not a generator. The block is written only when at least one signal is set, since without a signal there is nothing for it to explain.

Choosing What to Block

The right preset depends on what you want AI tools to do with your site. A few common situations:

You want to be cited, not trained on. This covers most blogs, publishers and documentation sites. Keep the default Training crawlers only preset: every training collector in the list gets a Disallow group, while AI search indexers, user-triggered fetchers and browsing agents stay allowed. Combine it with the default signals (training disallowed, search and AI input allowed). Your pages can still appear as sources in AI answers, which are a growing source of referral visits.

You want out completely. Choose All AI crawlers to block every bot in the catalogue, and consider setting AI input to Disallow as well. Be honest about the cost: AI search engines stop citing you, and assistants refuse to open your links even when a reader asks for them. For paywalled or licensed content that trade can be worth it. For a site that depends on being discovered, it rarely is.

You only want to state preferences. None (signals only) writes no bot groups at all, just the wildcard group with your Content-Signal line. This suits sites that are fine with being crawled but want their usage terms on record.

You want a specific mix. Click any bot chip to add or remove it and the preset switches to Custom by itself. Typical tweaks are the training preset plus PerplexityBot, or the training preset minus Google-Extended if you do not mind Gemini training on your pages.

You want to keep bots out of one folder. Set Disallow path to something like /drafts/ or /members/. Every selected bot is kept out of that path only, and the rest of the site stays open to it. Keep in mind that robots.txt is public, so listing a private folder also advertises that it exists. Anything truly private needs a login, not a robots rule.

Whichever option you pick, the wildcard group always ends with Allow: /, so regular search crawlers such as Googlebot and Bingbot are never blocked by this file.

Limits and Enforcement

robots.txt is a request, not a lock. RFC 9309, the standard behind the Robots Exclusion Protocol, says plainly that its rules are not a form of access authorization. Large operators such as OpenAI, Anthropic and Google publish their tokens and say they honour them, but nothing technically stops a scraper from ignoring the file or sending a browser-like user agent string.

Content-Signal is softer still. It records how you want content used, and support varies by operator. Google has said it does not act on Content-Signal, so for Gemini training the Google-Extended group is the control that matters.

A few more limits to keep in mind:

  • No retroactive effect. Blocking GPTBot today does not remove pages it collected earlier or models already trained on them.
  • Caching. Crawlers keep a copy of robots.txt for a while, and RFC 9309 asks them not to rely on a cached copy for more than 24 hours, so changes are not instant.
  • Tokens change. Operators launch and rename bots. The catalogue here is curated from operator documentation and community lists, so revisit your file every few months.

For real enforcement, pair robots.txt with rules at the server or CDN: user agent matching in your web server, firewall rules, or a bot management service that verifies crawlers by IP range or reverse DNS. A newer approach is cryptographic: the bot signs each request and the site checks the signature against a published key, which is the idea behind Web Bot Auth. A Web Bot Auth key generator produces that kind of key material and shows what a verified bot actually sends.

Before you upload, check the result against real URLs. A robots.txt tester lets you paste the file and see whether a given user agent may fetch a given path, which catches mistyped tokens and path errors before crawlers ever read the file.

robots.txt and llms.txt

The two files have similar names but do different jobs. robots.txt is an access policy: it tells crawlers which paths they may fetch. llms.txt is a proposed Markdown file at the site root that summarizes the site and points language models to its most useful pages. One says where bots may go; the other suggests what is worth reading.

robots.txt has no llms.txt directive. Crawlers generally skip fields they do not recognise, but a made-up line would be noise at best and could trip strict validators. So when you tick Point to llms.txt, the tool writes a comment instead, such as # LLM-readable site summary: https://example.com/llms.txt. People and tools reading the file can follow it, and no crawler will misread it.

Both pointers depend on the Site URL field. The address must start with http:// or https://, and the tool keeps only the origin, so https://example.com/blog/ becomes https://example.com. The Sitemap line is always written as Sitemap: https://example.com/sitemap.xml; if your sitemap lives elsewhere or uses a different file name, edit that line after downloading. Leave the field empty and both lines are skipped, even with their boxes ticked.

A sensible setup for a site that wants AI visibility on its own terms uses both files. robots.txt blocks training crawlers and sets the signals, while llms.txt gives AI search engines and assistants a clean map of the content you do want quoted. If you do not have one yet, an llms.txt generator can help you draft it. For classic search engines that support it, submitting new URLs with the IndexNow URL Submission Tool works alongside the Sitemap line rather than replacing it.

The ones we answer the most.

How do I block AI crawlers from my site?

Pick a preset in the tool, press Generate robots.txt, then download the file and upload it to the root of your domain. The default Training crawlers only preset suits most sites because it blocks training bots and keeps AI search visibility. If you already have a robots.txt, merge the new groups into it and keep a single User-agent: * group instead of overwriting your existing rules.

Where does the robots.txt file go?

It must sit at the root of the host, so it loads at https://example.com/robots.txt. Crawlers never look for it in subfolders. Each host needs its own file, so blog.example.com and shop.example.com are read separately.

Is this generator free, and do I need an account?

It is free and needs no sign up. The file is assembled in your browser from the options you choose, and nothing about your site is uploaded or stored. You can copy the result or download it as robots.txt.

What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

They are three OpenAI bots with three jobs. GPTBot collects pages that may be used to train models, OAI-SearchBot indexes pages so they can be cited in ChatGPT search, and ChatGPT-User fetches a page when a person asks ChatGPT to open it. Blocking only GPTBot keeps you out of training while the other two keep you visible, which is what the default preset does.

What is Content-Signal in robots.txt?

Content-Signal is a line that states how fetched content may be used, with three keys: search, ai-input and ai-train, each set to yes or no. It comes from Cloudflare's Content Signals Policy and sits in the User-agent: * group, so it addresses every crawler. The tool writes it as, for example, Content-Signal: search=yes, ai-input=yes, ai-train=no, and any key set to No preference is simply left out.

What does blocking Google-Extended actually do?

Google-Extended is a control token rather than a separate crawler: it tells Google not to use your pages for training Gemini models or grounding Gemini answers. Your pages are still fetched by Googlebot, so crawling, indexing and ranking in Google Search do not change. It also does not remove you from AI Overviews, which are part of Search.

Will blocking AI bots hurt my Google rankings?

Blocking training crawlers will not, because Google-Extended and the other training bots are separate from Googlebot, which this tool never blocks. Blocking AI search indexers does not touch Google rankings either, but it removes your pages from AI answer engines such as ChatGPT search and Perplexity and from the citations they show. That is why the default preset leaves those bots allowed.

Does robots.txt actually stop AI scraping?

No. It is a public request that well-behaved crawlers honour and others can ignore, and Google has said it does not act on Content-Signal. For real enforcement, add server-side or CDN rules that block or challenge unwanted user agents, ideally with bot verification. Blocking also does not remove anything collected before the rule existed. A robots.txt tester can at least confirm the file says what you intended.

Can I block AI bots from just one folder?

Yes. Type the folder into Disallow path, for example /articles/, and every bot you selected gets Disallow: /articles/ instead of Disallow: /. The path must start with a slash and applies to all blocked bots at once. If different bots need different paths, edit those groups by hand after downloading.

Why are the Sitemap and llms.txt lines missing from my file?

Both lines need a site URL. Ticking Add a Sitemap line or Point to llms.txt has no effect while the Site URL field is empty. Enter your full address, such as https://example.com, and generate again; the tool keeps only the origin and writes the /sitemap.xml line and the llms.txt comment against it.

Can I add a bot that is not in the list?

Not inside the tool, whose chips cover a curated catalogue of over 50 AI user agents. Download the file and add a group by hand: a User-agent: line with the bot's token exactly as its operator documents it, followed by a Disallow: line. Compliant crawlers match tokens without regard to case, but copying the documented spelling avoids surprises.

What is the difference between Content-Signal and Disallow?

Disallow asks a crawler not to fetch a path at all. Content-Signal lets the crawler fetch the page but states what the content may be used for afterwards, such as search allowed and training refused. They work at different layers, so the tool writes both: the signal in the wildcard group for every bot, and Disallow groups for the specific crawlers you want kept out entirely.

How is this different from Cloudflare's managed robots.txt?

Cloudflare's managed robots.txt is a dashboard setting for sites proxied through Cloudflare: it adds the Content Signals Policy and rules for AI training crawlers to your robots.txt automatically, with Cloudflare choosing the list. This generator works for any host, whether a static site, WordPress or your own server, and lets you choose each bot, the signal values and the disallow path yourself. The result is a plain file you own and can edit.