How to block AI crawlers in robots.txt
To block AI crawlers, add a User-agent group with Disallow: / for each AI bot you want to keep out, such as GPTBot, ClaudeBot and CCBot, to your robots.txt file. Leave Googlebot and the * group alone and your normal search visibility does not change.
Before you paste a list, decide what you are actually blocking. Some AI user agents collect training data. Others fetch pages to answer a search or a user's question in real time, and blocking those can remove you from AI answers that would have cited and linked you. robots.txt is also a request that well-behaved bots follow, not a lock.
Step by step
1. Sort the bots into training and retrieval
These are the user agents with published documentation. Check each vendor's crawler page before relying on a name, because vendors add and rename agents.
| User agent | Operator | What it is for |
|---|---|---|
GPTBot | OpenAI | Collecting content that may be used to train models |
OAI-SearchBot | OpenAI | Finding pages to show and link in ChatGPT search |
ChatGPT-User | OpenAI | Fetching a page because a user asked ChatGPT to |
ClaudeBot | Anthropic | Collecting content for model training |
PerplexityBot | Perplexity | Indexing pages for Perplexity answers and citations |
CCBot | Common Crawl | Building the open Common Crawl dataset, widely used for training |
Google-Extended | A control token, not a separate crawler: whether Google may use content for Gemini models | |
Applebot-Extended | Apple | A control token: whether Apple may use content Applebot crawled to train its models |
A common middle ground: block the training agents (GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended) and allow the search agents (OAI-SearchBot, PerplexityBot) so you can still be cited.
2. Understand what Google-Extended does not do
Blocking Google-Extended does not remove you from Google Search, and it does not keep you out of AI Overviews. AI Overviews are part of Search and use pages Googlebot crawls. Google says Google-Extended is not a ranking signal either.
If you want to limit how much of a page Google can show in AI Overviews, the controls are the same snippet controls used for regular results: nosnippet, max-snippet or data-nosnippet. Those also shrink your normal snippets, so use them with care.
3. Write the rules
A crawler obeys the most specific group that names it and ignores the * group. So each named bot gets its own complete group:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Several User-agent lines can share one set of rules, as above. To block only part of a site, for example premium articles, use Disallow: /members/ in the AI group instead of Disallow: /.
4. Save the file in WordPress
- Open Hydrogen SEO → Tools → Robots.txt Editor.
- Add your AI group above your existing rules. Hydrogen SEO adds the
Sitemap:line for you, so you do not need to duplicate it. - Click Save Changes.
The editor is locked if Settings → Reading → Discourage search engines is checked, or if a physical robots.txt exists in your web root. Fix either first. The robots.txt editor docs cover both.
5. Check the live file line by line
Click View Robots.txt or open yoursite.com/robots.txt. The current plugin version has a known issue where saving can remove the line breaks between rules, which leaves crawlers unable to parse them. Every directive must be on its own line. If they run together, fix and save again.
Then test a URL for each agent with the robots.txt tester: GPTBot on /blog/any-post/ should be blocked, Googlebot on the same URL allowed.
6. Watch your server logs
Your access logs show which user agents request pages and whether they come back after the change. Compliant bots stop requesting blocked paths after they refetch robots.txt, which usually happens within a day. A bot that keeps crawling blocked paths is either ignoring robots.txt or is not who it claims to be. For those, a firewall rule at your host or CDN is the real control.
What robots.txt cannot do
- It does not remove content already collected. Blocking today does not pull your pages out of existing training sets or datasets.
- It does not stop non-compliant scrapers. Anyone can ignore the file or fake a user agent.
- User-initiated fetches may be treated differently. OpenAI describes
ChatGPT-Useras acting for a user and has said robots.txt rules may not apply to it the same way. Read the vendor's current documentation if this matters to you. - llms.txt is not a blocking tool. It describes your site; it controls nothing. See adding llms.txt to WordPress.
Should you block at all?
If your business depends on being recommended in AI answers, blocking search and retrieval agents works against you. If your content is your product, for example paid research or original photography, blocking training agents is a reasonable default. Many sites do both: opt out of training, stay open to search. There is no ranking penalty in Google for either choice.
Common questions
Will blocking GPTBot remove me from ChatGPT search?
Not by itself. ChatGPT search results come from OAI-SearchBot, which you control separately. Block GPTBot for training and allow OAI-SearchBot if you want to stay citable.
Does blocking Google-Extended hurt my Google rankings?
No. Google states Google-Extended does not affect inclusion or ranking in Google Search, and it does not control AI Overviews.
How quickly do AI crawlers see the change?
Crawlers cache robots.txt and refetch it regularly, often within a day. Check your access logs to confirm.
Can I block AI bots with a plugin setting instead?
You still need robots.txt rules, or a firewall rule at your host or CDN for bots that ignore robots.txt.