How Detection Models Are Trained and Retrained
A detection model is a classifier trained on past traffic to score future traffic. The training loop looks conventional and behaves unusually, because both the labels and the underlying population are unstable.
Training starts with feature extraction
Raw observations are converted into a fixed feature vector: categorical fields encoded, timing distributions summarised into statistics, and sequences reduced to counts and transition patterns.
Feature engineering carries most of the weight in this domain, more than model architecture does. A well-chosen feature that captures a genuine inconsistency outperforms a sophisticated model given uninformative inputs.
Features are also chosen for interpretability, because every score must be explainable to whoever handles the complaint when a legitimate customer is affected.
Labels come from several imperfect sources
Some labels are near-certain: traffic that identified itself honestly as a crawler, or sessions that completed a verified authentication. These are reliable and cover only a fraction of traffic.
Others are inferred from outcomes, such as sessions later confirmed as fraudulent through a chargeback or an account recovery. These arrive weeks late and only for the subset that produced a visible consequence.
The remainder are labelled by existing rules, which means the model partly learns to reproduce the system it was meant to improve on. Recognising that limitation is more useful than pretending it away.
Validation must respect time
Splitting data randomly produces optimistic results, because sessions from the same operation appear on both sides of the split and the model effectively memorises them.
Realistic evaluation trains on one period and tests on a later one, which reproduces the actual deployment condition of scoring traffic that arrived after training finished.
The gap between random and time-based validation scores is often large, and it is the honest estimate of how the model will perform in production.
Retraining is continuous, not occasional
Browser releases change what ordinary traffic looks like on a schedule the site does not control, and a model trained before a major release can read the new normal as anomalous.
Traffic composition shifts for business reasons too. A marketing campaign that brings visitors from an unfamiliar region resembles coordinated activity in most feature spaces.
Regular retraining against fresh, freshly labelled traffic is therefore operational maintenance rather than improvement work, and systems that skip it degrade steadily instead of failing visibly.
Deployment is staged for a reason
New models are usually run alongside the existing one in scoring-only mode first, so their score distribution and disagreements can be compared against the incumbent before anything acts on them.
Rollout then proceeds by traffic share with the ability to revert quickly, because the failure mode is turning away real customers and that damage is not recoverable after the fact.