What is crawl budget?

Crawl budget is the number of URLs Googlebot can and wants to crawl on your site within a given time. Google describes it as the combination of two things: the crawl capacity limit, how much crawling your server can handle without slowing down, and crawl demand, how much Google wants to crawl based on how popular and fresh your URLs are.

For most sites, crawl budget is not a concern. Google's own guidance says it mainly matters for very large sites (roughly a million or more unique pages) or medium-sized sites (around ten thousand pages or more) whose content changes daily. If your site is smaller and new pages get indexed within a few days, you can skip optimizing it.

How it works

Capacity rises when your server responds quickly and falls when it slows down or returns 5xx errors. Demand rises for URLs that are popular, linked often, and change regularly; it falls for URLs Google considers stale or low value.

Crawl budget is wasted when bots spend their visits on URLs that should never be indexed. The usual culprits are faceted navigation creating endless parameter combinations, session IDs in URLs, soft 404s, long redirect chains, and duplicate content. Every request spent there is one not spent on your new or updated pages.

Example

An online store with 5,000 products and filters for color, size, price, and brand can expose millions of filter URLs such as /shoes/?color=red&size=9&sort=price_asc. Googlebot may spend most of its visits on those combinations while new products wait to be discovered.

The fix combines several tools: block pure-filter parameters in robots.txt, canonicalize remaining variants to the base category, keep only canonical URLs in the sitemap, and avoid linking to filter combinations that have no search value.

Search Console's Crawl stats report is the place to check. It shows total requests, breakdowns by response code and file type, and average response time.

In WordPress

A typical WordPress blog is far below the size where crawl budget matters. WooCommerce stores with layered navigation and sites with large archives are the exceptions. Useful steps: noindex or disallow internal search, avoid linking to filter parameters, fix redirect chains, and keep hosting fast. Hydrogen SEO's sitemap lists only indexable content, which keeps the discovery signal clean; see XML sitemaps.

Common mistakes

  • Using noindex to save crawl budget. Google still crawls noindexed pages to see the tag. robots.txt is the tool that stops crawling.
  • Worrying about it on a 200-page site. Content quality and internal links matter far more there.
  • Believing crawl-delay helps. Google ignores it.

Common questions

Does crawl budget affect rankings?

Not directly. It affects how quickly new and updated pages are discovered, which matters for large or fast-changing sites.

How do I increase my crawl budget?

Make the server faster and more reliable, remove low-value URLs from crawl paths, and publish content people link to. Google raises crawling when capacity and demand both grow.