What We Talk About When We Talk About Logs: Understanding the Effects of Dataset Quality on Endpoint Threat Detection Research
Jason Liu, Muhammad Adil Inam, Akul Goyal, Andy Riddle, Kim Westfall, Adam Bates
Abstract
Endpoint threat detection research hinges on the availability of worthwhile evaluation benchmarks, but experimenters' understanding of the contents of benchmark datasets is often limited. Typically, attention is only paid to the realism of attack behaviors, which comprises only a small percentage of the audit logs in the dataset, while other characteristics of the data are inscrutable and unknown. We propose a new set of questions for what to talk about when we talk about logs (i.e., datasets): What activities are in the dataset? We introduce a novel visualization that succinctly represents the totality of 100+ GB datasets by plotting the occurrence of provenance graph neighborhoods in a time series. How synthetic is the background activity? We perform autocorrelation analysis of provenance neighborhoods in the training split to identify process behaviors that occur at predictable intervals in the test split. Finally, How conspicuous is the malicious activity? We quantify the proportion of attack behaviors that are observed as benign neighborhoods in the training split as compared to previously-unseen attack neighborhoods. We then validate these questions by profiling the classification performance of state-of-the-art intrusion detection systems (R-CAID, FLASH, KAIROS, GNN) against a battery of public benchmark datasets (DARPA Transparent Computing and OpTC, ATLAS, ATLASv2). We demonstrate that synthetic background activities dramatically inflate True Negative Rates, while conspicuous malicious activities artificially boost True Positive Rates. Further, by explicitly controlling for these factors, we provide a more holistic picture of classifier performance. This work will elevate the dialogue surrounding threat detection datasets and will increase the rigor of threat detection experiments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8f85e4db-0889-4251-aefb-a4e61cd179ccCited by top-tier papers2
- CoLD: Collaborative Label Denoising Framework for Network Intrusion DetectionShuo Yang, Xinran Zheng, Jinze Li, Jinfeng Xu et al.NDSS 2026 · 1 citation
- DUPIN: Attack Learning Is Still Needed! Demonstrating Few-Shot after Unsupervised Pretraining Is A Nimble Forensics LearnerChanwoo Bae, Hailun Ding, Shiqing Ma, Xiangyu ZhangUSENIX Security 2026 · 1 citation
Related papers
- Kairos: Practical Intrusion Detection and Investigation using Whole-system ProvenanceZijun Cheng, Qiujian Lv, Jinyuan Liang, Yan Wang et al.S&P 2024 · 125 citations
- R-CAID: Embedding Root Cause Analysis within Provenance-based Intrusion DetectionAkul Goyal, Gang Wang, Adam BatesS&P 2024 · 40 citations
- SoK: History is a Vast Early Warning System: Auditing the Provenance of System IntrusionsMuhammad Adil Inam, Yinfang Chen, Akul Goyal, Jason Liu et al.S&P 2023
- Flash: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation LearningMati Ur Rehman, Hadi Ahmadi, Wajih Ul HassanS&P 2024 · 104 citations
- ATLAS: A Sequence-based Learning Approach for Attack InvestigationAbdulellah Alsaheel, Yuhong Nan, Shiqing Ma, Le Yu et al.USENIX Security 2021 · 256 citations
