From Texts to Rules: Generating Sigma Rules with Large Language Models from Cyber Threat Reports
Yongxin Cai, Jing Qiu, Qingming Li, Du Cheng, Lei Chen
摘要
Cyber Threat Reports (CTRs) deliver actionable intelligence essential for security systems detection rules. Large language models (LLMs) could serve as a bridge for CTRs-to-Rules translation through parsing and generation capabilities. However, the semantic disconnect and domain-specific constraints between high-level abstractions in CTRs and low-level machine semantics in rules fundamentally impede accurate detection rules generation. In this paper, we demonstrate that shell commands in CTRs can be effectively converted into Sigma detection rules for security systems. To this end, we propose SIGMERGE, an end-to-end framework that generates Sigma rules from texts of CTRs by constructing a semantic intermediate layer as a bridge. The SIGMERGE framework hierarchically organizes three modules by descending semantic levels: (1) The Information extraction module, high-level, utilizes a multi-subsequence algorithm and a fine-tuned domain-specific LLM, enabling accurate MITRE ATT&CK tactics, techniques, and procedures (TTPs) and command extractions; (2) The Attack description generation module, intermediate-level, employs preference optimization tuning with closed-loop self-validation to mitigate the semantic disconnect; (3) The Sigma rule generation module, machine-level, leverages a parameter-optimized retrieval algorithm to address domain-specific constraints. We constructed 7 datasets for training and conducted extensive experiments. To validate SIGMERGE, we evaluated it using 23 metrics against 16 baselines and 13 LLMs, and conducted 10 case studies integrated with real security systems to demonstrate both effectiveness and efficiency. Moreover, SIGMERGE has already contributed 4 novel Sigma rules to the official repository, all of which have been formally accepted.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Understanding the Reproducibility of Crowd-reported Security VulnerabilitiesDongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu 等USENIX Security 2018 · 被引用 138 次
- LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTIYuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici 等WWW 2025 · 被引用 38 次
- Parametric Retrieval Augmented GenerationWeihang Su, Yichen Tang, Qingyao Ai, Junxi Yan 等SIGIR 2025 · 被引用 25 次
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen 等WWW 2025 · 被引用 16 次
相关 Paper
- SoK: Automated TTP Extraction from CTI Reports - Are We There Yet?Marvin Büchel, Tommaso Paladini, Stefano Longari, Michele Carminati 等USENIX Security 2025
- A Knowledge Extraction Framework on Cyber Threat Reports with Enhanced Security ProfilesYongxin Cai, Jing Qiu, Fan Zhang, Qiang Li 等SIGIR 2025 · 被引用 6 次
- RulePilot: An LLM-Powered Agent for Security Rule GenerationHongtai Wang, Ming Xu, Yanpei Guo, Weili Han 等ICSE 2026 · 被引用 1 次
- Catch Me If You Can: Detector-Resistant Evasion via Semantics-Preserving Command Re-RealizationMuhammad Shoaib, Hare Sudhan Muthusamy, Tareq Alkhatib, Wajih Ul HassanS&P 2026 · 被引用 2 次
- RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command ExplainerJiangyi Deng, Xinfeng Li, Yanjiao Chen, Yijie Bai 等NDSS 2025
