What Makes a Signal Useful for Detection
Detection systems could collect almost anything a browser exposes, and most of it would be worthless. Deciding what to keep comes down to four properties, and a signal has to satisfy all of them to earn a place in a scoring model.
It has to separate the populations
The basic requirement is that the signal takes different values for automated and human traffic often enough to matter. A property that appears at the same rate in both carries no information whatever it costs to gather.
Separation is measured empirically against labelled traffic, and the measurement usually disappoints. Many intuitively suspicious properties turn out to be almost as common among ordinary visitors as among the traffic they were supposed to catch.
Weak separation is not automatically fatal, because many weak signals can combine into a strong one. It is fatal when the weakness is combined with any of the other failures below.
It has to resist cheap imitation
A signal that can be set to any value by editing one line in a script provides protection only against the least sophisticated traffic, and only until that line is edited.
Signals rooted in how software is built are far more durable, because changing them means changing the software rather than changing a setting. That imposes a real cost on whoever wants to change it.
The useful question is not whether a signal can be defeated but what defeating it costs. Detection is an economic exercise, and signals are chosen to raise cost rather than to be impossible to bypass.
It has to be stable for legitimate users
A signal that changes unpredictably for real visitors generates false positives regardless of how well it separates populations in the aggregate.
Stability is tested over time and across the diversity of real traffic, including older devices, unusual configurations, assistive technology and corporate networks that rewrite requests. These populations are small and they absorb false positives disproportionately.
Where a signal is stable for most users but erratic for a specific group, using it means accepting a known harm concentrated on that group, which is a decision that should be made deliberately rather than by default.
It has to be cheap to collect
Collection runs on every request, so cost compounds. A check that adds noticeable latency or measurable battery drain is paid for by every legitimate visitor to catch a small minority.
Some techniques are also detectable by the client, which changes behaviour and undermines the measurement. Heavy rendering work performed purely for identification falls into this category.
Cheapness includes maintenance cost. A signal that needs constant reference-data updates to stay meaningful consumes engineering attention indefinitely, and that has to be weighed against what it contributes.
Most candidates fail somewhere
The signals that survive are a small fraction of what browsers expose, and they tend to be ones nobody designed to be identifying. Incidental properties of implementation are harder to change than properties intended as configuration.
This is also why the useful set shrinks over time. As browser vendors reduce passive entropy, signals retire, and systems built on a fixed list rather than on a process for evaluating candidates decay without anyone noticing.