How Feedback Loops Corrupt a Detection Model
Detection systems act on their own predictions, which means they influence the data they will learn from next. That loop is the source of a failure mode that looks like improving accuracy while the system gets steadily worse.
Blocked traffic stops generating evidence
When a request is refused, the session ends there. Whatever that visitor would have done next is never observed, so the record contains the suspicion and none of the behaviour that would have tested it.
Traffic that was allowed produces a full history, including the outcomes used for labelling. The two groups are therefore observed under different conditions, which breaks the comparison the model depends on.
This is a censoring problem rather than a modelling one, and no amount of additional training data fixes it while the censoring continues.
Wrong decisions are recorded as correct
A legitimate user blocked in error usually leaves without complaining. No correction enters the system, and the block sits in the record as an unchallenged decision.
Retraining then treats that session as a confirmed example of the pattern it was flagged for, reinforcing precisely the mistake that produced it.
Over successive cycles the model becomes more confident about a boundary that was drawn incorrectly, and the confidence itself makes the error harder to find.
Metrics improve while performance degrades
Measured against its own labels the model appears to get better, because agreement between the model and labels derived from the model necessarily rises.
Meanwhile the real error rate on affected user segments grows invisibly, showing up in support volume and conversion rates rather than anywhere in the detection dashboard.
The disconnect between security metrics and business metrics during such a period is a reliable warning sign, and it is usually noticed first by people outside the security team.
Holdouts are the standard countermeasure
Reserving a small random share of traffic that is scored but never acted on preserves an uncensored sample, since those sessions play out fully regardless of their score.
That sample supports honest measurement of how the model would perform without enforcement, and it is the only place where false positives on blocked-looking traffic become visible.
The cost is accepting some abuse in the holdout, which is why holdouts are kept small and are usually excluded on the highest-risk endpoints.
Independent labels are the other half
Periodic manual review of a random sample provides labels the model did not produce, which is the only reliable way to detect that the boundary has drifted.
Complaint and appeal channels serve the same purpose in production, provided they are easy enough to use that affected users actually use them. An appeals process nobody can find produces the same silence as no process at all, and silence is what the loop feeds on.