Duplicate Content: What Actually Hurts Australian Business Sites (2026)
There is no duplicate content penalty. There are four specific duplication faults that do real damage to Australian business sites — and only one of them is the kind Google punishes.
Google does not have a duplicate content penalty. It has a canonicalisation process, which is a different thing with a different set of consequences. Google Search Central's documentation on consolidating duplicate URLs says the quiet part plainly: telling Google which version you prefer is optional, and "your site will likely do just fine without specifying a canonical preference." Where you don't nominate one, "Google will identify which version of the URL is objectively the best version to show to users in Search." That is a selection process, not a sanction. Google's old standalone duplicate content help page doesn't even exist as its own document anymore; the URL now redirects into that canonicalisation guidance.
This matters because the myth costs Australian businesses real money. Owners rewrite perfectly good service pages because a crawl tool flagged their Brunswick page and their Coburg page as near-identical. Agencies quote content rewrites for a problem that was actually a redirect misconfiguration. Meanwhile the four duplication faults that genuinely cost rankings sit untouched, because none of them show up in a similarity score.
What Google actually does with duplicates
When Google finds several URLs serving substantially the same thing, it groups them and picks one to represent the set. The others stay out of results, and the ranking signals — links, relevance, engagement — get consolidated onto the chosen version. Google states the preference explicitly for protocol variants: it "prefers HTTPS pages over equivalent HTTP pages as canonical, except when there are issues or conflicting signals," with an invalid certificate or insecure dependencies counted among those issues.
You can watch this happen. Google's page indexing report documentation defines two statuses worth knowing by heart. "Duplicate without user-selected canonical" means "this page is a duplicate of another page, although it doesn't indicate a preferred canonical page." And "Duplicate, Google chose different canonical than user" means "this page is marked as canonical for a set of pages, but Google thinks another URL makes a better canonical."
Neither is a punishment. Both are Google telling you it made a decision you didn't make for it. The damage comes from what it decided: the wrong page ranking, or a page you care about sitting silently outside the index while a near-twin takes its place.
The four cases that actually cause damage
1. Location-page boilerplate, which is the one Google really does punish
This is the exception to everything above. Fifteen suburb pages built from one template, with the suburb name swapped in the heading, the intro and the meta title, is not a canonicalisation problem. It's a spam policy problem.
Google's spam policies define doorway abuse to include "having multiple domain names or pages targeted at specific regions or cities that funnel users to one page." That description fits an enormous share of Australian trade and services sites, where the plumber in Preston has a page for every suburb inside a thirty-kilometre radius and all of them funnel to the same contact form. The related scaled content abuse policy covers the modern version of the same instinct: pages "generated for the primary purpose of manipulating search rankings and not helping users," which now explicitly includes using generative tools to spin out volume without adding value.
These are enforceable. Google's manual actions documentation lists "Thin content with little or no added value" as a manual action type, and names doorways among its examples alongside thin affiliate pages and scraped content. That's an actual human review resulting in an actual suppression — the closest thing to the penalty everyone worries about, applied to the duplication pattern almost nobody worries about.
The distinction Google draws is whether the page has a reason to exist beyond the search query. A location page that names the streets and building stock you actually work on, prices the travel, shows jobs done in that suburb and lists the local approvals or council quirks is a real page that happens to be geographic. A location page that is the Preston page with "Preston" replaced by "Northcote" is a doorway wearing a service page's clothes. We've written separately about what belongs on a location page that ranks and converts — the short version is that if you can't write 300 words about a suburb that would be wrong if you pasted them onto a different suburb's page, that page shouldn't exist yet.
2. The same page living at four addresses
The least glamorous fault and the most common. A single homepage is frequently reachable at http://yoursite.com.au, https://yoursite.com.au, http://www.yoursite.com.au and https://www.yoursite.com.au, and on some platforms also at /index.php or with and without a trailing slash. Add campaign tracking parameters and one page has a dozen addresses.
Google will usually sort this out on its own, and it will usually pick the version you'd have picked. Usually is doing real work in that sentence. The failure mode is when signals conflict: the sitemap lists the non-www version, the internal navigation links to www, the canonical tag points at HTTP because it was hardcoded in 2019, and half your inbound links from directory listings go to a fifth variant. Google resolves the conflict, but not necessarily the way you want, and any link equity that lands on a variant Google treats as separate rather than duplicate is equity doing nothing for you.
The fix is dull and permanent. Choose one canonical form, 301 every other variant to it at the server level, make internal links and the sitemap agree with that choice, and set the canonical tag to match. Google ranks the three canonicalisation methods by strength and puts permanent redirects at the top, ahead of rel="canonical" annotations and sitemap inclusion. Do the redirect and the rest becomes confirmation rather than argument.
Australian businesses hit a specific version of this when they own both .com.au and .com, or an NZ business runs .co.nz alongside .com. Two domains serving identical sites is the same problem at a larger scale, and the answer is the same: one canonical domain, the other redirected, not both live and hoping.
3. Staging leaked into the index
A staging or development copy of the site that Googlebot can reach is the duplication fault most likely to produce a genuinely alarming rankings graph, because it splits your site against itself at full strength. Every page exists twice, and the staging copy sometimes wins.
The instinct is to block it in robots.txt, which is not enough on its own. Google's own page indexing documentation is explicit that a robots.txt block "does not guarantee that the page won't be indexed through some other means" — if anything links to the staging URL, it can still surface. A noindex directive works, but only if Google can crawl the page to read it, which robots.txt then prevents. The two directives quietly cancel each other out, which is why so many staging sites are still in the index despite being "blocked".
The reliable answer is HTTP authentication on the staging environment. A password wall means Googlebot gets a 401 and there is nothing to index, no directive to misread, and no race between two files. It also stops a staging build carrying an indexing block into production, which is the mirror-image fault. If your rankings fell after a site move rather than a normal week, the migration-specific diagnostic covers the rest of that ground.
4. Syndication without a canonical
Publishing your content somewhere else — an industry association's news page, a franchisor's national blog, a supplier's partner directory, a press release distributed to twenty outlets — creates a copy on a domain that may well outrank you. Google picks one version to show, and its criteria are not "who published first." A high-authority industry body republishing your article will frequently be the version that survives the clustering.
The safeguards are straightforward and rarely requested. Ask the republisher for a rel="canonical" tag pointing back to your original, which tells Google to consolidate the signals onto your URL. Where that isn't possible, a noindex on their copy achieves the same outcome. Where neither is possible, a prominent link back to the original at least routes readers and some signal your way. The version worth avoiding is handing over the full text with nothing attached, then wondering why the association's page ranks for your words.
The reverse case is worth naming too. Filling out a website with manufacturer product descriptions, supplier blurbs or franchisor-issued copy puts you on the wrong side of Google's scraped content policy, which prohibits "republishing content from other sites without adding any original content or value."
What is not worth worrying about
Repeated boilerplate is not duplicate content in any sense that matters. Your footer, your address block, your compliance disclaimers, the qualifications paragraph under every practitioner bio, the shipping and returns text on every product page — none of that is a problem, and rewriting it fifteen ways to score better in a crawl tool makes the site worse, not better.
Similarity percentages from third-party SEO tools measure text overlap, not what Google does with the page. Two service pages at 70% similarity are fine if each one answers a question the other doesn't. Two service pages at 30% similarity can still both be doorways if neither one has a reason to exist. The percentage isn't the signal; the reason to exist is.
How to check your own site
Search Console's page indexing report is the only source of truth on this, because it's the only place that shows Google's decisions rather than a tool's guess. Look for the two duplicate statuses above — and if the rest of that report is what's worrying you, which Search Console statuses actually matter sorts the genuine faults from the ones that are Google confirming your own instructions. A short list of "duplicate, Google chose different canonical" entries usually means conflicting signals worth tidying. A long list of pages excluded as duplicates that you expected to rank means the underlying content isn't distinct enough, and no technical fix will change that. If pages that were previously indexed have since dropped out altogether, duplication is only one of the candidates, and the ranked list of causes is the faster route to the culprit.
Then check the obvious variants by hand. Load your homepage with and without www, over HTTP and HTTPS, and confirm each one lands on a single final address. Search Google for site:yourdomain.com.au and read what comes back for staging subdomains and stray variants. If nothing at all comes back, that's a different problem and this guide to why a site isn't showing up on Google in Australia is the better starting point.
Nothing outside Search Console can tell you which canonical Google picked. What the free site audit will do is rule out the platform-side conditions that make every one of these faults worse — whether robots.txt and your sitemap are letting Google crawl the site properly, whether the certificate is valid, and whether Core Web Vitals and page weight give Google a reason to prefer a competitor's version of the same information. It takes about a minute, and it eliminates the boring causes faster than reading a similarity report at 11pm.