Duplicate Content
Duplicate content is the same or near-identical content appearing at more than one URL, either on your own site or across different sites. Search engines must choose one version to rank, which splits signals and often surfaces the wrong page.
Why it matters for your rankings
There is no duplicate content penalty, despite two decades of rumour. What happens is dilution: links and relevance signals spread across several versions, so none of them ranks as well as one consolidated page would.
The largest source is technical rather than editorial. Parameter URLs, http and https, www and non-www, trailing slashes, pagination and printer-friendly versions can multiply a single article into a dozen crawlable URLs. Ecommerce adds another layer through manufacturer descriptions repeated across every retailer selling the same product. The fix is rarely rewriting. It is deciding which URL is authoritative and making that decision explicit through canonicals, redirects and consistent internal linking.
How to check it on your site
Find near-duplicates
Crawl with Screaming Frog and use Content then Near Duplicates. It reports pages above a similarity threshold you set.
Check protocol and host variants
Load your homepage as http, https, with and without www. All should redirect to one canonical version.
Search for scraped copies
Take a distinctive sentence from a key page and search it in quotes to find sites republishing your content.
Consolidate rather than rewrite
Pick the strongest URL, canonical or redirect the rest to it, and update internal links to point at the winner.