Toward Cybersecurity-Expert Small Language Models
Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sa, Ian Molloy, Yair Allouche
Abstract
Large language models (LLMs) are transforming everyday applications, yet they lag behind in specialized fields, such as cybersecurity, due to a lack of high-quality, domain-specific models and training datasets. To address this gap, we present CyberPal 2.0, a family of cybersecurity-expert small language models (SLMs) ranging from 4B–20B parameters. To train CyberPal 2.0, we generate an enriched chain-of-thought cybersecurity instruction dataset built with our data enrichment and formatting pipeline, SecKnowledge 2.0, which integrates expert-in-the-loop steering of reasoning formats alongside LLM-driven multi-step grounding, yielding higher-fidelity, task-grounded reasoning traces for security tasks. Across diverse cybersecurity benchmarks, CyberPal 2.0 consistently outperforms its baselines and matches or surpasses various open and closed-source frontier models , while remaining a fraction of their size. On core threat-investigation tasks, such as correlating vulnerabilities and bug tickets with weaknesses, our best 20B-parameter model outperforms GPT-4o, o1, o3-mini, and Sec-Gemini v1, ranking first, while our smallest 4B-parameter model ranks second. On core cyber threat intelligence knowledge tasks, our models outperform almost all tested frontier models, ranking second only to Sec-Gemini v1 . To foster reproducibility and practical adoption, we will release our models as open source.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- Instruction Tuning With Loss Over InstructionsZhengxiang Shi, Adam X. Yang, Bin Wu, Laurence Aitchison et al.NeurIPS 2024 · 55 citations
- CyberPal.AI: Empowering LLMs with Expert-Driven Cybersecurity InstructionsMatan Levi, Yair Allouche, Daniel Ohayon, Anton PuzanovAAAI 2025 · 17 citations
- CataractBot: An LLM-powered Expert-in-the-Loop Chatbot for Cataract PatientsPragnya Ramjee, Bhuvan Sachdeva, Satvik Golechha, Shreyas Kulkarni et al.UbiComp 2025 · 16 citations
Related papers
- The Digital Cybersecurity Expert: How Far Have We Come?Dawei Wang, Geng Zhou, Xianglong Li, Yu Bai et al.S&P 2025
- Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM TrainingYao-Ching Yu, Tsun-Han Chiang, Cheng-Wei Tsai, Chien-Ming Huang et al.EMNLP 2025 · 1 citation
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce et al.S&P 2024 · 167 citations
- Cyber-Zero: Training Cybersecurity Agents without RuntimeTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar et al.ICLR 2026 · 22 citations
- Benchmarking LLM-Assisted Blue Teaming via Standardized Threat HuntingYuqiao Meng, Luoxi Tang, Feiyang Yu, Xi Li et al.ICML 2026 · 6 citations
