Blog

Why Stability and Uniqueness Pull Against Each Other

Every fingerprinting design confronts the same tension. The properties that distinguish devices best are the properties most likely to change, and the properties that never change distinguish almost nothing.

Rare values identify, common values do not

Identification depends on rarity. An attribute whose value is shared by nearly everyone narrows the population barely at all, while one held by a handful of devices narrows it enormously.

The information a signal provides is a function of how surprising its value is, which is why detailed enumerable properties contribute so much more than coarse categorical ones.

That is also why the most identifying attributes are ones that reflect a specific machine's configuration rather than its general class, and configuration is exactly what changes.

The sharp attributes are the volatile ones

Installed fonts identify well because font sets accumulate through software installation and are close to unique on a well-used machine. They also change every time an application is installed.

Rendering output reflects the specific combination of hardware and driver, which is highly distinguishing. A driver update changes it, and driver updates arrive without the user's involvement.

Browser version is precise for a few weeks and then changes for the whole population at once. Precision and durability are inversely related across essentially the whole attribute set.

The stable attributes are shared

Screen resolution changes rarely, and it also takes one of a small number of standard values. Platform and architecture are stable and shared by millions.

An identifier built only from stable attributes therefore groups devices rather than distinguishing them. It survives indefinitely and identifies almost nobody.

This combination has its own uses, particularly for consistency checking, but it cannot support recognition of an individual device across visits.

Weighted matching is the practical compromise

Rather than choosing between the two, systems keep both and compare attribute by attribute with weights reflecting each attribute's volatility.

A mismatch on a volatile attribute counts for little, while a mismatch on a stable one counts heavily, because a device does not change architecture. The match threshold is then tuned against observed false-match and false-split rates.

This works, but it converts a clean identifier into a probability, and every downstream system has to be built to consume a probability rather than a key.

Both failure modes are costly

Setting the threshold loosely merges distinct devices, which corrupts counts and can attach one person's history to another. Setting it tightly splits a single device into many, which destroys continuity.

Neither error is symmetric in consequence, and which one to prefer depends entirely on the application. A fraud system usually prefers splitting; an analytics system usually prefers merging, and running one threshold for both purposes serves neither well.