What is indexing in SEO?

Indexing is the step where a search engine takes a page it has crawled, processes its content, and stores it in its index so it can be returned for relevant queries. A page that is not indexed cannot appear in search results, no matter how good it is.

Crawling and indexing are separate. Google crawls far more URLs than it indexes, and it makes an indexing decision for each URL based on quality, duplication, and your directives. Getting crawled is necessary; getting indexed is earned.

How it works

After fetching a page, Google renders it (running JavaScript where needed), extracts text, links, images, and structured data, and groups it with any duplicates. From each duplicate cluster it chooses one canonical URL to index. The others are recorded but not shown.

Pages can be excluded at this stage for several reasons:

  • A noindex directive.
  • Duplication, where another URL was chosen as canonical.
  • Low value, which shows up in Search Console as "Crawled - currently not indexed".
  • Not yet processed, shown as "Discovered - currently not indexed", often a sign Google does not see the page as a priority.
  • Errors: 404s, soft 404s, server errors, or redirects.

Indexing is not permanent. Pages can drop out when they go stale, when better duplicates appear, or after quality reassessments.

Example

To check a single URL, paste it into the URL Inspection tool in Google Search Console. It reports whether the page is on Google, which canonical Google chose, and when it was last crawled. You can request indexing from the same screen after fixing a problem.

For a site-wide view, the Indexing → Pages report groups every known URL by status and reason. A site:example.com search gives a rough idea, but its counts are estimates and should not be used for diagnosis.

In WordPress

The most common WordPress indexing problem is the Discourage search engines checkbox in Settings → Reading, which adds noindex everywhere. After that come thin archive pages (tags, date archives, attachment pages) that dilute the set of pages Google sees. SEO plugins let you noindex those types. In Hydrogen SEO that is the Robots screen, and its sitemap leaves noindexed types out automatically so the two signals agree; see robots meta settings.

Common mistakes

  • Expecting a sitemap to force indexing. It helps discovery only.
  • Requesting indexing over and over. It does not raise priority beyond the first request.
  • Treating every non-indexed URL as a problem. Redirects, noindexed pages, and parameter duplicates are supposed to be excluded.

Common questions

How long does it take Google to index a new page?

Anywhere from hours to weeks. Well-linked pages on sites Google crawls often are picked up fastest; pages with no internal links can wait a long time.

Why is my page crawled but not indexed?

Google fetched it and decided not to index it, usually because it judged the content thin, duplicative, or low in value compared with other pages. Improving the content and linking to it internally are the usual fixes.