R-CAID: Embedding Root Cause Analysis within Provenance-based Intrusion Detection
Akul Goyal, Gang Wang, Adam Bates
Abstract
In modern enterprise security, endpoint detection products fire an alert when process activity matches known attack behavior patterns. Human analysts then perform Root Cause Analysis (RCA) over event logs to determine if the alert is indicative of an actual attack. Data Provenance can help to automate RCA by representing event logs as a causal dependency graphs; in fact, researchers are now examining whether provenance-based anomaly detection should replace pattern-based detection altogether. Unfortunately, we observe that current approaches leverage off-the-shelf graph embedding techniques that are unable to associate events with their root causes. This shortcoming not only fails to capitalize on the RCA capabilities of provenance, but also leaves provenance-based IDS vulnerable to mimicry and evasion attacks.This work presents the design and implementation of R-CAID, a novel approach to incorporate RCA into provenance-based IDS. R-CAID precomputes each node’s root causes during graph construction, then directly links those nodes to their root causes during embedding. Further, R-CAID’s classification model is node/process-level, rather than graph/system-level, bringing it more in line with the precision of commercial systems. Under a passive adversary model, we find that R-CAID consistently outperforms baseline graph neural networks, sequence-based log IDS, and even a commercial endpoint detection system. Under a white-box active adversary model, R-CAID maintains a high level of performance (e.g., for DARPA Theia, 0.94 AUC adversarial down from 0.99 passive). R-CAID achieves this by associating each system entity with its immutable and unforgeable root causes, preventing adversaries from being able to masquerade as legitimate processes. This work is thus the first to demonstrate the promise of provenance-based IDS in a manner that avoids the pitfalls of mimicry and evasion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce8178eb-45ae-49ef-9dac-307b32498645Cited by top-tier papers10
- KnowHow: Automatically Applying High-Level CTI Knowledge for Interpretable and Accurate Provenance AnalysisYuhan Meng, Shaofei Li, Jiaping Gui, Peng Jiang et al.NDSS 2026 · 8 citations
- Entente: Cross-silo Intrusion Detection on Network Log Graphs with Federated LearningJiacen Xu, Chenang Li, Yu Zheng, Zhou LiNDSS 2026 · 3 citations
- Beyond Nodes vs. Edges: A Multi-View Fusion Framework for Provenance-Based Intrusion DetectionFan Yang, Binyan Xu, Di Tang, Kehuan ZhangS&P 2026 · 2 citations
- A Context Is Worth a Thousand Lies: Evading Intrusion Detectors via Intelligent Context DistortionMagdy Nasr, Vansh Rastogi, Azadeh TabibanBS&P 2026 · 1 citation
- Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection SystemsTristan Bilot, Baoxiang Jiang, Zefeng Li, Nour El Madhoun et al.USENIX Security 2025
Builds on20
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 1,823 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- HOLMES: Real-Time APT Detection through Correlation of Suspicious Information FlowsSadegh Momeni Milajerdi, Rigel Gjomemo, Birhanu Eshete, R. Sekar et al.S&P 2019 · 550 citations
- NoDoze: Combatting Threat Alert Fatigue with Automated Provenance TriageWajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen et al.NDSS 2019 · 411 citations
- Tactical Provenance Analysis for Endpoint Detection and Response SystemsWajih Ul Hassan, Adam Bates, Daniel MarinoS&P 2020 · 317 citations
Related papers
- Sometimes, You Aren't What You Do: Mimicry Attacks against Provenance Graph Host Intrusion Detection SystemsAkul Goyal, Xueyuan Han, Gang Wang, Adam BatesNDSS 2023
- vCause: Efficient and Verifiable Causality Analysis for Cloud-based Endpoint AuditingQiyang Song, Qihang Zhou, Xiaoqi Jia, Zhenyu Song et al.USENIX Security 2026
- SoK: History is a Vast Early Warning System: Auditing the Provenance of System IntrusionsMuhammad Adil Inam, Yinfang Chen, Akul Goyal, Jason Liu et al.S&P 2023
- ORTHRUS: Achieving High Quality of Attribution in Provenance-based Intrusion Detection SystemsBaoxiang Jiang, Tristan Bilot, Nour El Madhoun, Khaldoun Al Agha et al.USENIX Security 2025
- What We Talk About When We Talk About Logs: Understanding the Effects of Dataset Quality on Endpoint Threat Detection ResearchJason Liu, Muhammad Adil Inam, Akul Goyal, Andy Riddle et al.S&P 2025
