robots.txt tester

Paste your robots.txt, pick a crawler and enter a URL. The tester applies the same matching rules Google documents and shows whether the URL is allowed or blocked, which group applied and which line decided it. Nothing is sent to a server.

AllowedGooglebot may crawl /wp-admin/admin-ajax.php
  • Group used: User-agent: * (no group names this crawler, so it falls back to *)
  • Deciding rule, line 3: Allow: /wp-admin/admin-ajax.php

How the matching works

  1. Pick the group. A crawler follows only the group whose User-agent names it most specifically. Googlebot-Image uses a Googlebot-Image group if there is one, otherwise Googlebot, otherwise *. Groups are not combined across names, which is why a GPTBot group with one rule ignores everything under *.
  2. Pick the rule. Inside that group, the rule with the longest matching path wins. If an Allow and a Disallow match with the same length, Allow wins.
  3. Wildcards. * matches any run of characters and $ marks the end of the URL, so Disallow: /*.pdf$ blocks /guide.pdf but not /guide.pdf?download=1.

Paths are case-sensitive: Disallow: /Private does not block /private.

Blocked is not the same as removed

robots.txt controls crawling, not indexing. A blocked URL that other pages link to can still appear in Google with no description. To keep a page out of results, allow crawling and add a noindex tag instead. And never block CSS or JavaScript your pages need, or Google renders them broken.

Blocking AI crawlers

Most AI companies publish separate user-agents for training and for fetching pages to answer a question, for example GPTBot and OAI-SearchBot. Blocking one does not block the other. Google-Extended controls use of your content for Gemini training and does not affect Google Search. The AI crawler guide walks through the choices.

Common questions

Does Google follow crawl-delay?

No. Googlebot ignores the crawl-delay directive. Bing and some other crawlers honor it.

Where does WordPress keep robots.txt?

By default WordPress serves a virtual robots.txt generated on request, so there is no file on disk. A physical robots.txt in the site root overrides it, and SEO plugins, including Hydrogen SEO, let you edit the virtual one from wp-admin.

Why does my URL show as blocked in Search Console but allowed here?

Check that you pasted the live file from the same host and protocol, since www and non-www, or http and https, can serve different robots.txt files. Google also caches robots.txt for up to about a day.