Why Two Detection Vendors Disagree About One Visit
Running two detection products alongside each other and comparing their verdicts reliably produces disagreement on a substantial share of traffic. The disagreement is structural, and understanding why it happens is more useful than deciding which product is right.
They observe different things
A provider integrated at the network edge sees connection-level detail and the complete request as it arrived. One integrated as page script sees the browser environment and interaction but nothing below the application layer.
Neither view is complete, and the parts each misses are exactly the parts the other relies on. Two systems reasoning from disjoint evidence naturally reach different conclusions.
Integration point also determines timing. An edge product decides before the page loads, while a script-based product accumulates evidence over the session and can revise its view.
Their reference populations differ
Every provider calibrates against the traffic it sees across its own customers, and that mix varies enormously between a network serving global infrastructure and a vendor focused on financial services.
What counts as unusual is defined relative to that mix. A client configuration that is rare across one provider's customers may be commonplace across another's.
This means the same observation genuinely carries different information for each provider, and both can be correct about their own populations while disagreeing about a visit.
They are tuned for different error costs
A provider whose customers care most about content theft tunes towards catching more automation and accepts turning away some real visitors.
One whose customers care most about conversion tunes the other way, letting through traffic it suspects rather than risking friction on a purchase.
These are product decisions rather than technical ones, and they move the threshold far enough to produce disagreement on any session sitting near the middle of the range.
Score scales are not comparable
Providers use different ranges, different directions and different distributions. A middling value from one may correspond to a much stronger or weaker claim than the same number from another.
Comparing raw scores is therefore meaningless without calibrating both against a common set of sessions with known outcomes, which is work most teams never do.
Comparing the actions each would have taken is more informative, since the action encodes both the score and the threshold the provider considers appropriate.
Disagreement is itself a useful signal
Sessions where two systems agree are easy, and they are the bulk of traffic. Sessions where they disagree are concentrated in the genuinely ambiguous region.
Routing exactly those sessions to a stronger check or to human review spends the expensive resource where it changes outcomes, which is a better use of a second provider than treating either verdict as authoritative.