What is Googlebot?

Googlebot is the generic name for the two crawlers Google Search uses to download web pages: Googlebot Smartphone, which simulates a mobile user, and Googlebot Desktop. Because Google primarily indexes the mobile version of most sites, the smartphone crawler makes the large majority of requests.

Both crawlers answer to the same Googlebot token in robots.txt, so you can't allow one and block the other. They find new URLs mainly by following links from pages they already crawled, and they fetch CSS, JavaScript and images separately so Google can render the page the way a browser would.

Limits worth knowing

Google documents a few specifics that explain odd behavior:

  • Size cutoff. For Google Search, Googlebot fetches the first 2MB of a supported file type and the first 64MB of a PDF, measured uncompressed. Anything past the limit is ignored. Each CSS or JavaScript file referenced by a page is fetched separately under the same limit.
  • Crawl rate. For most sites, Googlebot shouldn't request pages more than once every few seconds on average, though short bursts can look faster. If your server struggles, you can reduce the rate.
  • Time zone. When crawling from US IP addresses, Googlebot's time zone is Pacific Time, which matters for pages that change content by time of day.

A WordPress page that inlines huge amounts of CSS, SVG or JSON can approach the HTML limit. Text beyond it isn't indexed.

Real Googlebot or an impostor?

Many scrapers set their user agent to Googlebot's to slip past firewalls. Google recommends checking the IP, not the user agent:

  1. Run a reverse DNS lookup on the IP from your logs, for example host 66.249.66.1.
  2. Confirm the hostname ends in googlebot.com, google.com or googleusercontent.com.
  3. Run a forward lookup on that hostname and confirm it returns the original IP.

For automated checks, match IPs against Google's published JSON lists of crawler IP ranges. Security plugins that block "fake Googlebots" usually do this check; ones that block by user agent alone can lock out the real crawler. For request-level evidence, see log file analysis.

Blocking it, and what that does

Match the tool to the goal. robots.txt stops crawling but doesn't guarantee a URL stays out of results; noindex keeps a page out of the index but requires Google to crawl it to see the tag; password protection keeps everyone out. Blocking Googlebot affects Google Search, Discover, Images, Video and News. For per-crawler meta tags, see the googlebot meta tag.

Other Google crawlers

Googlebot is only one of Google's crawlers. Googlebot Image and Googlebot Video have their own limits, and Google runs other agents such as GoogleOther for non-Search work. Google-Extended isn't a separate crawler at all: it's a robots.txt token that controls whether content Google crawls can be used for its Gemini models.

Common questions

Can I block Googlebot Desktop but allow Googlebot Smartphone?

Not with robots.txt. Both crawlers use the same Googlebot token, so a rule applies to both.

Why does Googlebot ignore part of my page?

Googlebot fetches only the first 2MB of an HTML file for Google Search. Content past that point isn't considered for indexing.

Is every request with a Googlebot user agent from Google?

No. The user agent is often faked. Verify with a reverse DNS lookup and a forward lookup, or match the IP against Google's published ranges.