USENIX Security2026Top-tier venue
From Texts to Rules: Generating Sigma Rules with Large Language Models from Cyber Threat Reports
Yongxin Cai, Jing Qiu, Qingming Li, Du Cheng, Lei Chen
Abstract
Cyber Threat Reports (CTRs) deliver actionable intelligence essential for security systems detection rules. Large language models (LLMs) could serve as a bridge for CTRs-to-Rules translation through parsing and generation capabilities. However, the semantic disconnect and domain-specific constraints between high-level abstractions in CTRs and low-level machine semantics in rules fundamentally impede accurate detection rules generation. In this paper, we demonstrate that shell commands in CTRs can be effectively converted into Sigma detection rules for security systems. To this end, we propose SIGMERGE, an end-to-end framework that generates Sigma rules from texts of CTRs by constructing a semantic intermediate layer as a bridge. The SIGMERGE framework hierarchically organizes three modules by descending semantic levels: (1) The Information extraction module, high-level, utilizes a multi-subsequence algorithm and a fine-tuned domain-specific LLM, enabling accurate MITRE ATT&CK tactics, techniques, and procedures (TTPs) and command extractions; (2) The Attack description generation module, intermediate-level, employs preference optimization tuning with closed-loop self-validation to mitigate the semantic disconnect; (3) The Sigma rule generation module, machine-level, leverages a parameter-optimized retrieval algorithm to address domain-specific constraints. We constructed 7 datasets for training and conducted extensive experiments. To validate SIGMERGE, we evaluated it using 23 metrics against 16 baselines and 13 LLMs, and conducted 10 case studies integrated with real security systems to demonstrate both effectiveness and efficiency. Moreover, SIGMERGE has already contributed 4 novel Sigma rules to the official repository, all of which have been formally accepted.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cacb22d2-d061-4b2d-8cde-74b7aecbf828Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Understanding the Reproducibility of Crowd-reported Security VulnerabilitiesDongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu et al.USENIX Security 2018 · 138 citations
- LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTIYuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici et al.WWW 2025 · 38 citations
- Parametric Retrieval Augmented GenerationWeihang Su, Yichen Tang, Qingyao Ai, Junxi Yan et al.SIGIR 2025 · 25 citations
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen et al.WWW 2025 · 16 citations
Related papers
- SoK: Automated TTP Extraction from CTI Reports - Are We There Yet?Marvin Büchel, Tommaso Paladini, Stefano Longari, Michele Carminati et al.USENIX Security 2025
- A Knowledge Extraction Framework on Cyber Threat Reports with Enhanced Security ProfilesYongxin Cai, Jing Qiu, Fan Zhang, Qiang Li et al.SIGIR 2025 · 6 citations
- RulePilot: An LLM-Powered Agent for Security Rule GenerationHongtai Wang, Ming Xu, Yanpei Guo, Weili Han et al.ICSE 2026 · 1 citation
- Catch Me If You Can: Detector-Resistant Evasion via Semantics-Preserving Command Re-RealizationMuhammad Shoaib, Hare Sudhan Muthusamy, Tareq Alkhatib, Wajih Ul HassanS&P 2026 · 2 citations
- RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command ExplainerJiangyi Deng, Xinfeng Li, Yanjiao Chen, Yijie Bai et al.NDSS 2025
