How Risk Scores Combine Independent Signals
A risk score is produced by combining many weak indicators into one number. The combination step is where most of the engineering difficulty sits, and adding signals up is not what happens.
Signals are weighted by how much they discriminate
A signal earns weight by separating the two populations. If a property appears at similar rates in automated and human traffic, it contributes nothing regardless of how easy it is to collect.
Discrimination is measured against labelled traffic, which is why the quality of labels bounds the quality of the weights. Weak labels produce confidently wrong weights.
Weights are not fixed once. A signal that discriminated well loses value as the traffic it identified adapts, so weighting is maintenance rather than a setup step.
Correlated signals must not be counted twice
Many indicators measure overlapping things. Timezone, language preference and network location all speak to geography, and treating them as independent triples the apparent evidence.
Naive combination therefore produces overconfident scores. A single underlying oddity gets amplified into what looks like several independent confirmations.
Handling this means either modelling the correlation explicitly or grouping related indicators into a single composite feature. Either way, the redundancy has to be identified first.
Context changes what a signal is worth
The same observation means different things on different endpoints. An unusual client on a public page is unremarkable; the same client posting credentials repeatedly is not.
Time of day, request rate and account history all modulate weights. Systems that score a request in isolation discard the context that would resolve the ambiguity.
This is why detection is usually session-scoped rather than request-scoped. A sequence of actions carries far more information than any single one of them.
Models drift as populations change
Browser releases change what normal traffic looks like on a schedule nobody controls. A model trained before a release can read the new normal as anomalous.
Seasonal and campaign-driven shifts do the same thing. A sudden influx of genuine visitors from an unfamiliar network resembles coordinated activity in most feature spaces.
Continuous recalibration against fresh traffic is the only durable answer. Systems left untouched degrade steadily rather than failing visibly.
Explainability constrains model choice
Teams need to know why a request scored as it did, both to debug complaints and to satisfy internal review. That requirement rules out some model families.
Many production systems therefore keep an interpretable scoring layer over more complex components, so that every decision reduces to a small set of named contributing factors.