ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
Yiran Wu, Mauricio Velazco, Andrew Zhao, Manuel Luján, Srisuma Movva, Yogesh Roy, Quang Nguyen, Roberto Rodriguez, Qingyun Wu, Michael Albada, Julia Kiseleva, Anand Mudgerikar
Abstract
We present ExCyTIn-Bench, the first benchmark to Evaluate an LLM agent X on the task of Cyber Threat Investigation through security questions derived from investigation graphs. Real-world security analysts must sift through a large number of heterogeneous security logs, follow multi-hop chains of evidence to investigate threats. With the developments of LLMs, building LLM-based agents for automatic threat investigation is a promising direction. We construct a benchmark from a controlled Azure tenant including a SQL environment covering 57 log tables from Microsoft Sentinel and related services, and 7542 generated questions. We leverage security logs extracted with expert-crafted detection logic to build threat investigation graphs, and then generate questions with LLMs using paired nodes on the graph, taking the start node as background context and the end node as answer. Anchoring each question to these explicit nodes and edges not only provides automatic, explainable ground truth answers but also makes the pipeline reusable and readily extensible to new logs. Our comprehensive experiments on the test set with different models confirm the difficulty of the task: the best model so far can achieve a reward of 0.606, leaving much headroom for future research. The code is available at SecRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3772c2d1-443f-47a3-91de-385a255acdebBuilds on9
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin et al.AAAI 2024 · 484 citations
- Tactical Provenance Analysis for Endpoint Detection and Response SystemsWajih Ul Hassan, Adam Bates, Daniel MarinoS&P 2020 · 317 citations
- SLEUTH: Real-time Attack Scenario Reconstruction from COTS Audit DataMd Nahid Hossain, Sadegh M. Milajerdi, Junao Wang, Birhanu Eshete et al.USENIX Security 2017 · 291 citations
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 209 citations
Related papers
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- DRBench: A Realistic Benchmark for Enterprise Deep ResearchAmirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox et al.ICLR 2026 · 18 citations
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr et al.ICML 2025
- Benchmarking LLM-Assisted Blue Teaming via Standardized Threat HuntingYuqiao Meng, Luoxi Tang, Feiyang Yu, Xi Li et al.ICML 2026 · 6 citations
- InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight GenerationGaurav Sahu, Abhay Puri, Juan A. Rodríguez, Amirhossein Abaskohi et al.ICLR 2025
