Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training
Yao-Ching Yu, Tsun-Han Chiang, Cheng-Wei Tsai, Chien-Ming Huang, Wen-Kwang Tsao
摘要
Large Language Models (LLMs) have shown remarkable advancements in specialized fields such as finance, law, and medicine. However, in cybersecurity, we have noticed a lack of opensource datasets, with a particular lack of highquality cybersecurity pretraining corpora, even though much research indicates that LLMs acquire their knowledge during pretraining. To address this, we present a comprehensive suite of datasets covering all major training stages, including pretraining, instruction fine-tuning, and reasoning distillation with cybersecurityspecific self-reflection data. Extensive ablation studies demonstrate their effectiveness on public cybersecurity benchmarks. In particular, continued pre-training on our dataset yields a 15.9% improvement in the aggregate score, while reasoning distillation leads to a 15.8% gain in security certification (CISSP). We will release all datasets and trained cybersecurity LLMs under the ODC-BY and MIT licenses to encourage further research in the community. 1 * Primary Contributor. † Equal Contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet TextXiming Lu, David Acuna, Jaehun Jung, Jian Hu 等ICML 2026 · 被引用 6 次
- Careful Queries, Credible Results: Teaching RAG Models Advanced Web Search Tools with Reinforcement LearningYuqin Dai, Shuo Yang, Guoqing Wang, Yong Deng 等AAAI 2026 · 被引用 5 次
- RedSage: A Cybersecurity Generalist LLMNaufal Suryanto, Muzammal Naseer, Pengfei Li, Syed Talal Wasim 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang 等ICML 2024 · 被引用 605 次
相关 Paper
- Toward Cybersecurity-Expert Small Language ModelsMatan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sa 等ICML 2026 · 被引用 7 次
- CyberPal.AI: Empowering LLMs with Expert-Driven Cybersecurity InstructionsMatan Levi, Yair Allouche, Daniel Ohayon, Anton PuzanovAAAI 2025 · 被引用 17 次
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza 等ICLR 2026 · 被引用 22 次
- Rewriting Pre-Training Data Boosts LLM Performance in Math and CodeKazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura 等ICLR 2026 · 被引用 21 次
- MAmmoTH2: Scaling Instructions from the WebXiang Yue, Tianyu Zheng, Ge Zhang, Wenhu ChenNeurIPS 2024 · 被引用 176 次
