A Knowledge Extraction Framework on Cyber Threat Reports with Enhanced Security Profiles
Yongxin Cai, Jing Qiu, Fan Zhang, Qiang Li, Lei Chen
Abstract
Knowledge extraction on Cyber Threat Reports (CTRs) is critical for attack investigation and defenses. The granularity and the usability of the knowledge are key issues: the former is determined by entity recognition on CTRs, whereas the latter mainly depends on proper relation extraction. Nevertheless, in the state-of-the-art entity recognition methods on CTRs using span representation, the local semantics of behavior are not considered and the sequential features of entity labels within behavior descriptions are not utilized. Besides, domain-specific definitions/forms of the relation types and knowledge representations are also crucial for effective utilization of knowledge. In this paper, we propose a novel knowledge extraction framework on CTRs to address the above concerns. The framework is formed by the Enhanced Security Profiles (ESP) that can be directly utilized by security detection devices. In the ESP framework, we propose 3 modules to facilitate fine-grained and accurate knowledge extractions: (1) The entity recognition module utilizes a label-aware subsequence autoregressive algorithm to integrate local semantic and label sequence features, enabling accurate identification of cybersecurity entities; (2) The relation extraction module employs LLM-based strategies with shared partition representations to enhance semantic understanding and domain relevance; and (3) The security profile generation module leverages Chain-of-Thought reasoning and In-Context Learning to produce machine-readable rules executable in security detection systems. Extensive experiments on 6 datasets demonstrate that the ESP framework largely outperform the state-of-the-art solutions e.g., the Micro-Fl scores on entity recognition and relation extraction are at least 1.54% and 13.12% better, respectively. Our code can be found in https://github.com/YxinMiracle/ESP.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e26e6847-352e-42a0-ae68-73a7fa74470dCited by top-tier papers1
Ask how each one uses itRelated papers
- KnowHow: Automatically Applying High-Level CTI Knowledge for Interpretable and Accurate Provenance AnalysisYuhan Meng, Shaofei Li, Jiaping Gui, Peng Jiang et al.NDSS 2026 · 8 citations
- SoK: Automated TTP Extraction from CTI Reports - Are We There Yet?Marvin Büchel, Tommaso Paladini, Stefano Longari, Michele Carminati et al.USENIX Security 2025
- LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTIYuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici et al.WWW 2025 · 38 citations
- Toward Cybersecurity-Expert Small Language ModelsMatan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sa et al.ICML 2026 · 7 citations
- REACT: Residual-Adaptive Contextual Tuning for Fast Model Adaptation in Threat DetectionJiayun Zhang, Junshen Xu, Bugra Can, Yi FanWWW 2025 · 3 citations
