Choose language

Robots.txt Tester

Test whether URLs are allowed or blocked for any crawler, with the exact rule and line that decides each result. Follows RFC 9309.

Robots.txt TesterHow it works ↓
Paste the file as served at /robots.txt. Evaluated in your browser, nothing is fetched or uploaded.
A product token such as Googlebot, GPTBot or ClaudeBot, or a full User-Agent string
One per line. Full URLs are reduced to their path and query.

Follows RFC 9309, including its * wildcard and $ end anchor: the crawler's own group wins over the * group, the longest matching rule wins, and a tie goes to Allow. Real crawlers may still differ in edge cases.

Mehmet Demiray Published Updated
Share

How robots.txt Matching Works

A robots.txt file is a list of groups. Each group opens with one or more User-agent lines and continues with the Allow and Disallow rules that apply to those crawlers. A crawler reading the file does not combine every group it finds. It picks one set of rules and ignores the rest.

The tester follows the same selection order RFC 9309 describes:

  1. The crawler's own token. If any group names the product token, for example GPTBot, that group applies. The comparison is case-insensitive, so gptbot and GPTBot are the same agent.
  2. A parent token. If a hyphenated token has no group of its own, the tester falls back to the longest parent that does. Googlebot-Image uses the Googlebot group when there is no Googlebot-Image group, which mirrors how Google treats its sub-crawlers.
  3. The * group. Only when neither of the above exists does the wildcard group apply.
  4. Nothing. With no matching group and no * group, every URL is allowed.

The summary card names which of these happened, so you can see at a glance whether your crawler hit its own rules or fell through to *.

The point that trips people up most is that a specific group replaces the * group instead of adding to it. Say your file has Disallow: /private/ under User-agent: *, followed by a separate User-agent: GPTBot group containing only Allow: /. GPTBot may crawl /private/, because the rules under * simply do not exist for it. If a rule should apply to a named crawler, repeat it inside that crawler's group.

Two smaller details also come from the spec. Several groups naming the same agent are merged into one, so Googlebot rules split across the file behave as if they were written together. And consecutive User-agent lines before any rule share one group, which is how you give GPTBot and ClaudeBot identical rules without writing them twice.

The crawler field accepts a bare product token or a full User-Agent string. With a full string the tester tries to extract the token first. If the Group applied row shows the * group when you expected a named one, enter the bare token instead; that always works.

Longest Match Wins: Allow vs Disallow

Once the group is chosen, the tester compares each URL path against every Allow and Disallow rule in it. The order of the lines does not matter. Length does: among all rules whose pattern matches the path, the one with the longest pattern decides. If an Allow and a Disallow of equal length both match, Allow wins. If no rule matches at all, the URL is allowed.

Take a group with four rules: Disallow: /shop/, Allow: /shop/sale/, Disallow: /*.pdf$ and Disallow: /*?sort=. Here is how some paths come out:

Path Matching rules Verdict
/shop/cart Disallow: /shop/ Blocked
/shop/sale/shoes Disallow: /shop/, Allow: /shop/sale/ Allowed, longer pattern
/files/guide.pdf Disallow: /*.pdf$ Blocked
/files/guide.pdf?v=2 none Allowed by default
/shop?sort=price Disallow: /*?sort= Blocked
/blog/ none Allowed by default

* matches any run of characters, including none and including slashes. A $ at the end of a pattern means the path must end exactly there, which is why /*.pdf$ blocks the PDF but not the same PDF with a query string attached. Without $, every rule is a prefix match: Disallow: /admin also blocks /administrator and /admin-login. Write /admin/ if you only mean the directory.

A Disallow: line with nothing after it blocks nothing. It is the traditional way of saying everything is allowed, and the tester skips it when matching. Disallow: / is the opposite case, the shortest pattern that matches every path, so any more specific Allow beats it.

For every URL the results table shows the winning rule with its line number, for example "Line 4: Allow /shop/sale/". When it says "No rule matched (allowed by default)", nothing in the applied group covers that path. Ties are rare, but they do happen with pairs like Allow: /page and Disallow: /page, and the tester resolves them toward Allow, the least restrictive choice the spec asks for.

Testing AI Crawlers like GPTBot and ClaudeBot

AI companies publish their crawler tokens so site owners can address them in robots.txt: GPTBot for OpenAI, ClaudeBot for Anthropic, PerplexityBot, CCBot and others, plus Google-Extended, a control token Google reads to decide whether content fetched by its regular crawlers may be used for its AI models. Testing any of them works exactly like testing Googlebot. Type the token in the crawler field, paste the file and list the URLs you care about.

Group selection matters a lot here. Many sites start with a broad User-agent: * group and then add a short block for each AI crawler. Because a named group replaces *, an AI crawler with its own group ignores every rule written for everyone else, including the Disallow lines protecting admin or checkout paths. Running the same URL list for GPTBot, then ClaudeBot, then a crawler you never mentioned is a fast way to confirm each one sees what you intended.

Some files now also carry a Content-Signal line, part of the Content Signals Policy proposed by Cloudflare. It states how content may be used after it has been fetched, with three keys:

  • search: building a search index and showing links and short excerpts in results
  • ai-input: feeding content into AI answers at query time, such as retrieval or grounding
  • ai-train: training or fine-tuning AI models

Each key takes yes or no, for example search=yes, ai-input=yes, ai-train=no.

The tester reports the Content-Signal value from the group applied to your crawler in the summary card, and it never changes a verdict. A signal is a preference; Allow and Disallow decide whether fetching is permitted. A crawler can be allowed to fetch a page whose signal says ai-train=no, and whether that preference is respected is up to the crawler's operator. Since the line lives inside a group, a crawler with its own group will not see a signal placed under * unless you repeat it there.

If you are writing these rules from scratch, an AI crawler robots.txt generator can produce the per-crawler groups, and pasting its output here lets you test them before anything goes live.

Common robots.txt Mistakes

The parser flags lines that crawlers will ignore or misread and lists them under Parser warnings with their line numbers. These come up most often:

  • Rules before any User-agent line. An Allow or Disallow at the top of the file belongs to no group, so crawlers drop it. A Content-Signal line placed above the first group is ignored for the same reason.
  • Paths that do not start with / or *. Disallow: private/ is not a valid pattern. Write /private/.
  • Crawl-delay. It is not part of RFC 9309 and Google ignores it. Bing does read it, so the tester shows a warning rather than treating it as broken.
  • Vendor directives. Host and Clean-param are Yandex extensions. Noindex, Nofollow, Request-rate and Visit-time are flagged the same way. Google stopped honouring Noindex in robots.txt on September 1, 2019, so it does nothing there.
  • Typos and unknown fields. Dissalow: /tmp/ is reported as an unknown directive. Crawlers skip it silently, which leaves /tmp/ open.
  • Broken lines, such as a line with no colon or a User-agent: with no value.

Two mistakes produce no warning at all, because the syntax is valid.

First, paths are case-sensitive. Disallow: /Private/ does not block /private/, and the tester will show that path as allowed. Directive names are not case-sensitive, so disallow: and Disallow: behave the same, but the paths after them are compared character by character.

Second, blocking is not the same as removing from search. robots.txt controls crawling. A blocked page that other sites link to can still show up in results, usually without a description, because the crawler was never allowed to read it. To keep a page out of the index, it has to stay crawlable and carry a noindex meta tag or HTTP header.

One more trap comes from prefix matching. Disallow: /blog also blocks /blog-archive and /blogroll. When you test the path you mean to block, add a few neighbouring paths to the list as well and check that only the intended ones come back Blocked.

What This Tester Does Not Do

The tester evaluates the text you paste. It never fetches a live robots.txt, so a draft that has never been deployed, such as the output of an AI crawler robots.txt generator, can be checked the same way as a production file. It also ignores the domain of the URLs you list. A full URL is reduced to its path and query string, and the rules in your pasted file are applied to that whatever host the URL names. To test a live site, open its /robots.txt in a browser tab and copy the contents in.

It does not check whether a page is indexed, ranked or even reachable. Allowed means the rules permit fetching, nothing more. HTTP behaviour is outside its scope too. A real crawler that gets a 404 for robots.txt treats everything as allowed, and repeated server errors can make it treat the whole site as off limits for a while. Pasted text has no status code, so none of that is simulated.

Path normalisation is limited. Patterns are compared exactly as written, and the tester does not decode or re-encode percent escapes in your rules. When you paste a full URL, the browser's URL parser may encode characters such as spaces or accented letters, so a rule written with the raw character can miss the encoded path. For URLs with non-ASCII characters, write the rule in percent-encoded form, which is how crawlers request them.

The matching follows RFC 9309, including the * and $ special characters, and Google, Bing and the main AI crawlers share that core logic. Individual crawlers can still differ in edge cases, and some ignore robots.txt entirely. robots.txt cannot prove who a crawler is either, since any client can send a GPTBot user agent.

Everything runs in your browser. Nothing is uploaded, there is no account to create and the tool is free, so internal or unreleased files stay on your device. The per-URL results can be exported as CSV or as an image for a ticket or an audit report.

The ones we answer the most.

How do I test if a URL is blocked by robots.txt?

Paste the contents of the robots.txt into the first box, enter the crawler you want to check, such as Googlebot or GPTBot, and list the URLs or paths one per line. Press Test. Each URL gets an Allowed or Blocked verdict with the deciding rule and its line number, and the summary shows which User-agent group was applied. If no rule matched, the table says so and the URL counts as allowed.

Do I need an account or a live website to use this tester?

No. The tester is free, needs no sign up and works on any text you paste, including a draft that has never been deployed. Nothing is fetched or uploaded. The file is parsed and matched in your browser, so private or unreleased rules stay on your device.

Can I paste full URLs instead of paths?

Yes. A full address starting with https:// is reduced to its path and query before matching, so a blog post link with a tracking parameter is tested as something like /blog/post?utm=1. The domain is ignored, so the rules from your pasted file apply whatever host the URL names. A line without a leading slash is treated as a path with / added in front.

What happens when an Allow and a Disallow rule both match the same URL?

The rule with the longer pattern wins, no matter which comes first in the file. With Disallow: /shop/ and Allow: /shop/sale/, the path /shop/sale/shoes is allowed because the Allow pattern is longer. If both patterns have the same length, Allow wins. The results table names the winning rule, so you can see exactly why.

What do the * and $ characters mean in robots.txt rules?

* matches any run of characters, including none, and $ at the end of a pattern means the path must end there. Disallow: /*.pdf$ blocks /files/guide.pdf but not /files/guide.pdf?v=2. Neither was part of the original robots.txt convention, but RFC 9309 now describes both, and Google, Bing and the major AI crawlers honour them, so the tester does too.

What is the Content-Signal line in robots.txt?

Content-Signal is a line inside a User-agent group that states how fetched content may be used. It has three keys, search, ai-input and ai-train, each set to yes or no. It expresses a preference and blocks nothing technically. The tester shows the value from the group applied to your crawler, but it never changes an Allowed or Blocked verdict because of it.

My crawler isn't named in the robots.txt. Which rules apply to it?

It falls back to the User-agent: * group if the file has one, and the summary shows "Rules for * (all crawlers)". If there is no * group either, the result reads "No matching group; everything is allowed". A hyphenated token such as Googlebot-News first tries its parent token before falling back to *.

Why does Googlebot-Image follow the Googlebot rules?

When a hyphenated token like Googlebot-Image has no group of its own, the tester uses the group of its longest parent token, here Googlebot, the same way Google handles its sub-crawlers. The summary labels this as a parent token match. Add a User-agent: Googlebot-Image group if that crawler should follow different rules.

Why is Crawl-delay flagged as a warning?

Crawl-delay is not part of RFC 9309, and Google ignores it entirely. Bing still reads it, so keeping the line is not wrong, but it will not slow Googlebot down. The tester lists it under Parser warnings with its line number and does not use it in any verdict.

Does blocking a URL in robots.txt remove it from Google search results?

No. robots.txt controls crawling, not indexing. A blocked URL can still be indexed when other pages link to it, usually shown without a description because Google could not read the page. To remove a page from results, let it be crawled and add a noindex meta tag or HTTP header, since a blocked crawler never sees that tag.

How is this different from the robots.txt report in Google Search Console?

Search Console shows the robots.txt files Google last fetched from your verified site, along with fetch status and errors, and it only reflects Google's crawlers. This tester takes any pasted text, so you can check a draft before deploying it, try any token including GPTBot or ClaudeBot, and see the deciding line for each URL. It needs no site verification, but it cannot tell you what Google actually fetched.