Blog

How Session Replay Data Feeds Detection

Session replay was built to let product teams watch how people use a page. The event stream it produces is also a rich detection input, and that dual use is where most of the difficulty lies.

What replay actually captures

Replay does not record video. It records a serialised snapshot of the page structure followed by a stream of mutations, pointer positions, scroll offsets, input events and viewport changes.

Playback reconstructs the session by replaying those mutations against the snapshot. The fidelity is high because the reconstruction uses the same rendering engine that produced the original.

The volume is substantial, since pointer and mutation events are frequent. Sampling and compression are applied aggressively to keep the stream manageable.

The same stream answers detection questions

Everything a behavioural model needs is present: interaction timing, movement dynamics, event ordering and the relationship between actions and page state.

Because capture happens for analytics anyway, using it for detection adds no collection cost. This is why the two functions frequently share infrastructure inside a product.

It also gives detection something rare: full context for a flagged session, so an analyst can see what actually happened rather than inferring it from a score.

Field masking is the critical control

Replay would otherwise capture whatever users type, including credentials, payment details and personal information entered into forms.

Implementations therefore mask inputs by default and require explicit opt-in to record specific fields. Masking that operates before data leaves the device is the only version that provides a real guarantee.

Misconfiguration here is a recurring source of incidents, usually because a custom component was not recognised by the masking rules and fell through as plain text.

Purpose limitation becomes hard to maintain

Data collected for product analytics being used for security decisions is a change of purpose, and consent frameworks in several jurisdictions treat those as distinct.

The distinction matters practically as well. Analytics data is typically retained longer and shared more widely internally than security data, and the combination inherits the loosest handling of the two.

Separating the pipelines, even at some duplication cost, is the cleaner arrangement and the one that survives review.

Retention drives most of the risk

Replay archives accumulate detailed records of individual sessions, which is a meaningful concentration of behavioural data about identifiable people.

Short retention windows, aggressive sampling and deriving features early so that raw streams can be discarded all reduce that exposure without much loss of usefulness, since both analytics and detection consume aggregates far more often than individual sessions.