Why Deduplication Errors Skew Analytics
Every unique-visitor figure rests on a matching decision, and that decision is never perfect. The two ways it fails push counts in opposite directions, which is why reported audience figures are estimates presented as facts.
Splitting inflates counts
When one person's visits fail to match, each visit is counted as a new visitor. Cleared storage, a browser update or a change of network can all cause this.
Inflation is systematic rather than random, because the causes correlate with user behaviour. Privacy-conscious users split most often and are therefore overrepresented in visitor counts.
The effect is largest over long reporting periods, since more opportunities to split accumulate. Monthly figures are more inflated than daily ones for the same underlying audience.
Merging deflates counts and corrupts segments
When distinct people match to the same identity, they are counted once. Shared devices, standardised corporate fleets and collision-prone environments all cause this.
The count problem is the lesser issue. The merged record combines two people's behaviour into one profile, so anything derived from behaviour becomes an average of unrelated individuals.
Segmentation built on such records describes a population that does not exist, and the artefacts look like genuine findings because nothing in the data marks them as merged.
The two errors do not cancel
It is tempting to assume inflation and deflation offset each other, but they affect different populations and different metrics.
Splitting affects users who manage their privacy actively; merging affects users on shared or standardised devices. These are largely disjoint groups with different characteristics.
The net effect on any given metric therefore depends on the audience composition, and it changes as that composition shifts, which makes trends unreliable in a way that is hard to see.
Cross-device matching adds another layer
Attempting to unify a person's phone, laptop and tablet into one identity introduces a second matching decision on top of the first, with its own error rates.
Deterministic matching through authentication is reliable and covers only logged-in users, so it is applied to a subset and extrapolated to the rest.
Probabilistic matching covers everyone and errs frequently, particularly by merging household members whose devices share a network and usage patterns.
Ranges are more honest than point estimates
Because the error is structural, reporting a single number implies a precision the method cannot deliver. Comparisons between periods inherit the same uncertainty.
Teams that measure their own match rates against authenticated traffic can quantify the error and report accordingly, which usually reveals that small period-over-period movements are noise rather than signal.