What is robots.txt?
robots.txt is a plain-text file served at https://yourdomain.com/robots.txt that tells crawlers which parts of a site they may fetch. Well-behaved bots, including Googlebot and Bingbot, request it before crawling and follow its rules. The format dates from 1994 and was formalized as an internet standard (RFC 9309) in 2022.
The single most important fact about robots.txt: it controls crawling, not indexing. A disallowed URL can still appear in search results, usually as a bare URL with no description, if other pages link to it. To keep a page out of results, use noindex instead.
How it works
The file is a list of groups. Each group starts with one or more User-agent lines naming the crawler it applies to, followed by Allow and Disallow rules matching URL paths. * as a user agent means all crawlers. When rules conflict, Google uses the most specific (longest) matching rule. A Sitemap: line can point crawlers at your XML sitemap, and it can appear anywhere in the file.
Each host and protocol needs its own file: blog.example.com does not inherit rules from example.com. Google ignores the crawl-delay directive, although some other crawlers honor it. Rules are public, so never use robots.txt to hide sensitive paths; it effectively lists them for anyone who looks.
Many site owners now also use robots.txt to address AI crawlers by user agent, such as GPTBot or Google-Extended. Honoring those rules is up to each operator.
Example
A typical WordPress robots.txt:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Sitemap: https://example.com/sitemap_index.xml
This keeps bots out of the admin area and internal search results while leaving the AJAX endpoint some themes rely on reachable.
In WordPress
WordPress serves a virtual robots.txt when no physical file exists in the site root. A physical file, if present, takes priority over anything a plugin generates. Hydrogen SEO includes a robots.txt editor under Tools, adds your sitemap line automatically, and has a reset option back to a safe default. The editor is locked while Discourage search engines is checked in Settings → Reading.
Common mistakes
- A stray
Disallow: /left over from staging, which blocks the whole site. - Blocking CSS and JavaScript. Google renders pages and needs those files to see the layout.
- Disallowing pages you want deindexed. The crawler can no longer see the noindex tag.
- Assuming it is private. Anyone can read the file.
Common questions
Does every site need a robots.txt file?
No, but it is recommended. Without one, crawlers assume everything is allowed, and a request for the missing file returns a 404, which is harmless but leaves you no place to declare a sitemap.
Can robots.txt remove a page from Google?
No. It only stops crawling. Use a noindex tag, keep the page crawlable, and let Google recrawl it.