USENIX Security2026Top-tier venue
DUPIN: Attack Learning Is Still Needed! Demonstrating Few-Shot after Unsupervised Pretraining Is A Nimble Forensics Learner
Chanwoo Bae, Hailun Ding, Shiqing Ma, Xiangyu Zhang
Abstract
Advanced persistent threats (APTs) pose a significant challenge in cybersecurity, involving staged and prolonged operations that often remain undetected until postmortem indicators, such as sabotage or financial loss, emerge. In consequence, human analysts are facing a needle-in-a-haystack challenge among vast accumulation of daily audit logs. However, leveraging a learning system for attack forensics is limited due to the scarcity of attack data, as attacks occur infrequently. In addition, since malicious behaviors are usually embedded in massive benign ones, they are very hard to label by humans. Thus, the recent approaches leverage self-supervised learning methods, where models rely solely on benign data and perform outlier detection. However, these methods struggle with the increasing complexity and dynamics of large-scale audit logs, often resulting in non-trivial false positives. Therefore, we propose a novel approach to learning-based attack forensics called DUPIN. First, DUPIN performs unsupervised pre-training on an enormous amount of audit events in the form of provenance graphs. It then proceeds to a few-shot learning stage, leveraging a small number of labeled attack examples to fine-tune its detection capabilities. We pretrain DUPIN on up to 38 - 52 days of audit logs (7.3TB total) and evaluate it against various baselines on 25 APT campaigns across four different data sources, facilitating the scalable evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on28
- HOLMES: Real-Time APT Detection through Correlation of Suspicious Information FlowsSadegh Momeni Milajerdi, Rigel Gjomemo, Birhanu Eshete, R. Sekar et al.S&P 2019 · 550 citations
- Pure Transformers are Powerful Graph LearnersJinwoo Kim, Dat Nguyen, Seonwoo Min, Sungjun Cho et al.NeurIPS 2022 · 311 citations
- Acing the IOC Game: Toward Automatic Discovery and Analysis of Open-Source Cyber Threat IntelligenceXiaojing Liao, Kan Yuan, XiaoFeng Wang, Zhou Li et al.CCS 2016 · 308 citations
- GraphFormers: GNN-nested Transformers for Representation Learning on Textual GraphJunhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li et al.NeurIPS 2021 · 262 citations
- ATLAS: A Sequence-based Learning Approach for Attack InvestigationAbdulellah Alsaheel, Yuhong Nan, Shiqing Ma, Le Yu et al.USENIX Security 2021 · 256 citations
Related papers
- Slot: Provenance-Driven APT Detection through Graph Reinforcement LearningWei Qiao, Yebo Feng, Teng Li, Zhuo Ma et al.CCS 2025 · 1 citation
- MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation LearningZian Jia, Yun Xiong, Yuhong Nan, Yao Zhang et al.USENIX Security 2024 · 92 citations
- TAPAS: An Efficient Online APT Detection with Task-guided Process Provenance Graph Segmentation and AnalysisBo Zhang, Yansong Gao, Changlong Yu, Boyu Kuang et al.USENIX Security 2025
- Unicorn: Runtime Provenance-Based Detector for Advanced Persistent ThreatsXueyuan Han, Thomas F. J.-M. Pasquier, Adam Bates, James Mickens et al.NDSS 2020
- TREC: APT Tactic / Technique Recognition via Few-Shot Provenance Subgraph LearningMingqi Lv, Hongzhe Gao, Xuebo Qiu, Tieming Chen et al.CCS 2024 · 18 citations
