Why Labelled Bot Data Is Hard to Obtain
Supervised detection needs examples known to be automated and examples known to be human. Obtaining either with confidence is genuinely difficult, and the difficulty shapes what detection systems can and cannot achieve.
Confirmation is rare and delayed
Truly confirmed automated traffic usually announces itself. Well-behaved crawlers identify themselves and can be verified through their operator's published infrastructure, which produces clean labels for the least interesting category.
Traffic that was actively hiding is confirmed only when it causes a consequence somebody investigates: a disputed transaction, a compromised account, a content scrape that surfaces elsewhere.
Those confirmations arrive weeks or months after the session, by which time the operation has changed. The labels describe a version of the problem that no longer exists.
Only harmful automation gets confirmed
Automated traffic that causes no visible harm is never investigated and therefore never labelled. Scraping that nobody notices and inventory checking that costs nothing sit permanently in the unlabelled pool.
The labelled set is consequently biased towards a specific slice of automation rather than being a sample of it. A model trained on that slice learns the characteristics of harmful operations, not of automation in general.
Whether that bias matters depends on the goal. It is acceptable when the aim is preventing harm and misleading when the aim is measuring how much traffic is automated.
Human labels are not certain either
A completed purchase or a successful authentication suggests a human, but neither is proof. Automation operating with stolen credentials produces exactly these outcomes.
Challenge completion is weaker still, since challenge-solving services exist and a solved challenge only demonstrates that somebody somewhere solved it.
The strongest human evidence tends to be long behavioural history: an account used consistently over years in ways that would be expensive to simulate. That is unavailable for new visitors, who are precisely where uncertainty is highest.
Rule-generated labels create circularity
Most labelled data comes from what existing rules flagged, which means a model trained on it learns to agree with those rules, including their mistakes.
Traffic the rules never flagged is labelled human by default, so any automation the current system misses is taught to the new model as an example of legitimate traffic.
Breaking the circle requires labelling a random sample independently, which is slow and expensive manual work. Teams that do it usually find their rules were wrong in ways nobody suspected.
Unsupervised methods fill the gap
Because labels are scarce, much production detection relies on anomaly detection and clustering, which need no labels and instead look for traffic that differs from the bulk or that clusters unnaturally tightly.
These methods find coordination well and intent not at all, so they are typically used to surface candidates for human review rather than to make decisions directly. The scarce labels are then spent on evaluating those candidates rather than on training from scratch.