What is index bloat?

Index bloat is when a search engine indexes many URLs from your site that have no business ranking: empty tag archives, internal search results, filter and sort combinations, attachment pages, and old test content. The pages you care about are still there, but they share the index with hundreds or thousands of near-empty ones.

It is not a penalty and it has no official definition from Google. It is a symptom. A bloated index usually means your site is producing URLs faster than it produces value, and that pattern costs you crawl attention and makes the site as a whole look thinner than it is.

How to spot it

Compare two numbers: how many pages you would actually want in search results, and how many Google reports as indexed in Search Console's Pages report. If you have 150 posts and pages but 2,400 indexed URLs, something is generating the extra 2,250.

To find the source, open the indexed list in the Pages report and look for patterns in the URLs:

  • /tag/ or /category/ archives with one or two posts each
  • /?s= internal search result pages
  • ?orderby=, ?filter_color= and other parameter URLs from faceted navigation
  • /attachment/ pages that show a single image
  • /page/37/ deep pagination on archives nobody browses

Why it matters

A few stray URLs do no harm. Bloat becomes a problem at scale, for three reasons:

  1. Crawl attention goes to the wrong URLs. On larger sites, Googlebot spends visits on junk instead of your new and updated content. See crawl budget.
  2. Quality signals are sitewide in part. Google has said in the past that large amounts of unhelpful content can affect how the rest of a site performs. Thousands of thin pages are a poor look.
  3. Diagnostics get noisy. Real problems hide inside a report full of URLs you never meant to publish.

How to fix it

Pick the tool that matches each URL type:

URL typeFix
Useful to visitors, not to searchers (thin tags, search results)noindex, keep links followable
Duplicate of another page (parameters, print views)canonical to the main URL
Should not exist at all (attachment pages, test posts)301 redirect or delete and let it 404
Infinite crawl spaces (filter combinations)block in robots.txt

Then remove those URLs from your XML sitemap and stop linking to them internally. In WordPress, Hydrogen SEO lets you noindex whole post types and taxonomies from one screen; the guides on noindexing tag and category archives and redirecting attachment pages cover the two most common sources. For low-value posts, prune thin content.

Expect the indexed count to fall slowly. Google has to recrawl each URL to see the new signal.

Common questions

Is index bloat a Google penalty?

No. It is a description of a site state, not a manual action or algorithm. The risk is indirect: wasted crawling and a larger share of weak pages.

Should I block bloated URLs in robots.txt?

Only for URLs you never want crawled. If a URL is already indexed and you block it, Google cannot see a noindex tag on it, so noindex first and block later if needed.

How many indexed pages is too many?

There is no fixed number. The useful test is whether the indexed count roughly matches the number of pages you would be happy to see in results.