What is duplicate content?

Duplicate content is content that appears, identically or almost identically, at more than one URL. It can be internal (several URLs on your site serving the same page) or external (your article republished on another site).

Duplicate content is not a penalty in itself. Google has said plainly that it does not penalize sites for ordinary duplication. What it does is group the duplicates, pick one as the canonical, and show only that one. The cost is losing control: Google may pick the wrong URL, and links and other signals get spread across versions. Actual penalties are reserved for deliberate scraping and spam.

How it works

Most duplication is technical and unintentional:

  • Protocol and host variants: http and https, www and non-www.
  • Trailing slashes and case: /Guide/ and /guide.
  • Parameters: tracking codes, session IDs, sort orders, filters.
  • Print or AMP versions, and paginated comments.
  • CMS archives: the same post listed at category, tag, and date URLs (these are usually excerpts and less of an issue).
  • Product variants with separate URLs but the same description.
  • Syndication and scraped copies on other domains.

The tools to consolidate are canonical tags (keep both URLs, credit one), 301 redirects (remove the duplicate), and consistent internal linking to the preferred URL. For pages that should exist but not rank, noindex.

Example

A product is reachable at:

code
https://example.com/grinders/burr-grinder-x/
https://example.com/grinders/burr-grinder-x/?color=black
https://example.com/sale/burr-grinder-x/

All three serve the same product. The fix: every version carries <link rel="canonical" href="https://example.com/grinders/burr-grinder-x/">, internal links use that URL, and only it appears in the sitemap. If the /sale/ path is not needed after the sale, 301 it to the main URL.

In WordPress

WordPress creates predictable duplicates: attachment pages, tag and date archives, ?replytocom comment links on older setups, and category base variations. SEO plugins output self-referencing canonicals and let you noindex thin archive types. Hydrogen SEO does both; its site audit also finds duplicate titles and descriptions across pages, which often point to duplicate content. See canonical URLs and site audits.

Common mistakes

  • Canonicalizing every page to the home page.
  • Using robots.txt to block duplicates, which stops Google seeing the canonical tag.
  • Copying manufacturer product descriptions word for word across hundreds of products and expecting them to rank.

Common questions

Is there a duplicate content penalty?

Not for normal duplication. Google filters duplicates and shows one version. Penalties apply only to deliberate, manipulative copying such as scraped content.

How much similarity counts as duplicate?

Google does not publish a percentage. Pages whose main content is the same, apart from small differences like a color name or a tracking parameter, are treated as duplicates.