How Fingerprint Hashes Are Actually Computed
A browser fingerprint is usually presented as a short hexadecimal string, which makes it look like a fixed property of the device. It is the output of a pipeline, and each stage of that pipeline shapes what the value means.
Collection produces a set of attribute values
The first stage queries a fixed list of sources: rendering results, enumerated capabilities, screen and locale properties, and details of how requests are formed.
Each source returns a value of some kind, and the collection is deliberately ordered. The order matters because the hash is computed over a sequence, not an unordered set.
Sources that fail or are unavailable still occupy a slot, typically with a placeholder. A missing value is itself informative, so it is recorded rather than skipped.
Normalisation decides what counts as the same
Raw values are noisy. Floating-point rendering results vary slightly between runs, and lists come back in inconsistent orders on some platforms.
Normalisation rounds, sorts and canonicalises those values so that a device produces the same input twice. Getting this wrong makes fingerprints unstable for reasons unrelated to the device.
How aggressively you normalise is a direct trade-off. Coarse rounding improves stability but merges devices that genuinely differ, which reduces the value of the result.
Hashing turns the vector into a short label
The normalised vector is serialised and passed through a hash function, producing a fixed-length string. Any general-purpose hash works, because the goal is compactness rather than security.
The hash is not reversible in any useful sense, but it is not private either. Anyone with the same collection code and the same device produces the same output.
Storing the hash instead of the raw vector is common but costly. It discards the per-attribute detail that later comparison would need.
A single changed attribute changes the whole hash
Hash functions are designed so that a small input change produces an unrelated output. That property is useful for integrity checking and unhelpful for identification.
A font installed, a browser updated, a monitor changed: any of these flips one slot and the resulting label bears no resemblance to the previous one.
Systems that key on the hash alone therefore see a stream of new devices. The churn is an artefact of the representation, not of real device turnover.
Why fuzzy matching sits on top
Mature systems keep the attribute vector and compare vectors directly, counting how many slots agree. Two vectors differing in one field are treated as a probable match.
Attributes are weighted by how likely they are to change, so a browser version mismatch counts for less than a mismatch in rendering output. The hash then becomes a fast lookup key rather than the identity itself.