DrSec: Flexible Distributed Representations for Efficient Endpoint Security
Mahmood Sharif, Pubali Datta, Andy Riddle, Kim Westfall, Adam Bates, Vijay Ganti, Matthew Lentz, David Ott
Abstract
The increasing complexity of attacks has given rise to varied security applications tackling profound tasks, ranging from alert triage to attack reconstruction. Yet, security products, such as Endpoint Detection and Response, bring together applications that are developed in isolation, trigger many false positives, miss actual attacks, and produce limited labels useful in supervised learning schemes. To address these challenges, we propose DrSec—a system employing self-supervised learning to pre-train foundation language models (LMs) that ingest event-sequence data and emit distributed representations for processes. Once pre-trained, the LMs can be adapted to solve different downstream tasks with limited to no supervision, helping unify the currently fractured application ecosystem. We trained DrSec with two LM types on a real-world dataset containing ∼91M processes and ∼2.55B events, and tested it in three application domains. We found that DrSec enables accurate, unsupervised process identification; outperforms leading methods on alert triage to reduce alert fatigue (e.g., 75.11% vs. ≤64.31% precision-recall area under curve); and accurately learns expert-developed rules, allowing tuning incident detectors to control false positives and negatives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 464421c4-d6e9-4673-8e01-48dc65ddea74Cited by top-tier papers3
- Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection SystemsTristan Bilot, Baoxiang Jiang, Zefeng Li, Nour El Madhoun et al.USENIX Security 2025
- WatchLog: Efficient and Interpretable Event Reasoning for Endpoint Detection and Response Logs with Multimodal LLMsHongyi Zhou, Jianfeng Pan, Min Peng, Shaomang Huang et al.ICML 2026
- HyperGLLM: An Efficient Framework for Endpoint Threat Detection via Hypergraph-Enhanced Large Language ModelsHongyi Zhou, Jianfeng Pan, Min Peng, Shaomang Huang et al.AAAI 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski et al.NDSS 2019 · 826 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 447 citations
- NoDoze: Combatting Threat Alert Fatigue with Automated Provenance TriageWajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen et al.NDSS 2019 · 411 citations
Related papers
- ANTEATER: A Filter-then-Scrutinize Architecture for End-to-End Attack InvestigationYiming Ren, Haoqiang Wang, Linghao Li, Haoyang Chen et al.SIGMOD 2026
- Incident Response Planning Using a Lightweight Large Language Model with Reduced HallucinationKim Hammar, Tansu Alpcan, Emil C. LupuNDSS 2026 · 16 citations
- Beyond Raw Bytes: Towards Large Malware Language ModelsLuke Kurlandski, Harel Berger, Yin Pan, Matthew WrightNDSS 2026 · 5 citations
- ATLAS: A Sequence-based Learning Approach for Attack InvestigationAbdulellah Alsaheel, Yuhong Nan, Shiqing Ma, Le Yu et al.USENIX Security 2021 · 256 citations
- GARNET: GoT-Based Alert Reduction and Narrative Event TracingYiru Gong, Song Liu, Changzhi Zhao, Junrong Liu et al.AAAI 2026
