LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
Yuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici, Asaf Shabtai
Abstract
As the number and sophistication of cyber attacks have increased, threat hunting has become a critical aspect of active security, enabling proactive detection and mitigation of threats before they cause significant harm. Open-source cyber threat intelligence (OS-CTI) is a valuable resource for threat hunters, however, it often comes in unstructured formats that require further manual analysis. Previous studies aimed at automating OSCTI analysis are limited since (1) they failed to provide actionable outputs, (2) they did not take advantage of images present in OSCTI sources, and (3) they focused on on-premises environments, overlooking the growing importance of cloud environments. To address these gaps, we propose LLMCloudHunter, a novel framework that leverages large language models (LLMs) to automatically generate generic-signature detection rule candidates from textual and visual OSCTI data. We evaluated the quality of the rules generated by the proposed framework using 12 annotated real-world cloud threat reports. The results show that our framework achieved a precision of 92% and recall of 98% for the task of accurately extracting API calls made by the threat actor and a precision of 99% with a recall of 98% for IoCs. Additionally, 99.18% of the generated detection rule candidates were successfully compiled and converted into Splunk queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cbec8eb-0f6d-48f5-a8b0-d6626f304f68Cited by top-tier papers4
- RulePilot: An LLM-Powered Agent for Security Rule GenerationHongtai Wang, Ming Xu, Yanpei Guo, Weili Han et al.ICSE 2026 · 1 citation
- From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration TestingLanxiao Huang, Daksh Dave, Tyler Cody, Peter A. Beling et al.EMNLP 2025 · 1 citation
- A Unified Framework for Rule Learning: Integrating Commonsense Knowledge from LLMs with Structured Knowledge from Knowledge GraphsQirui Hao, Kewei Cheng, Tongze Zhang, Hongyuan Liu et al.WWW 2026
- From Texts to Rules: Generating Sigma Rules with Large Language Models from Cyber Threat ReportsYongxin Cai, Jing Qiu, Qingming Li, Du Cheng et al.USENIX Security 2026
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- CASIE: Extracting Cybersecurity Event Information from TextTaneeya Satyapanich, Francis Ferraro, Tim FininAAAI 2020 · 148 citations
- Enabling Efficient Cyber Threat Hunting With Cyber Threat IntelligencePeng Gao, Fei Shao, Xiaoyuan Liu, Xusheng Xiao et al.ICDE 2021 · 124 citations
Related papers
- APT-CGLP: Advanced Persistent Threat Hunting via Contrastive Graph-Language Pre-TrainingXuebo Qiu, Mingqi Lv, Yimei Zhang, Tieming Chen et al.KDD 2026
- Benchmarking LLM-Assisted Blue Teaming via Standardized Threat HuntingYuqiao Meng, Luoxi Tang, Feiyang Yu, Xi Li et al.ICML 2026 · 6 citations
- SoK: Automated TTP Extraction from CTI Reports - Are We There Yet?Marvin Büchel, Tommaso Paladini, Stefano Longari, Michele Carminati et al.USENIX Security 2025
- Incident Response Planning Using a Lightweight Large Language Model with Reduced HallucinationKim Hammar, Tansu Alpcan, Emil C. LupuNDSS 2026 · 16 citations
- TaintP2X: Detecting Taint-Style Prompt-to-Anything Injection Vulnerabilities in LLM-Integrated ApplicationsJunjie He, Shenao Wang, Yanjie Zhao, Xinyi Hou et al.ICSE 2026
