CmdCaliper: A Semantic-Aware Command-Line Embedding Model and Dataset for Security Research
Sian-Yao Huang, Cheng-Lin Yang, Che-Yu Lin, Chun-Ying Huang
摘要
This research addresses command-line embedding in cybersecurity, a field obstructed by the lack of comprehensive datasets due to privacy and regulation concerns. We propose the first dataset of similar command lines, named CyPHER 1 , for training and unbiased evaluation. The training set is generated using a set of large language models (LLMs) comprising 28,520 similar command-line pairs. Our testing dataset consists of 2,807 similar command-line pairs sourced from authentic command-line data. In addition, we propose a command-line embedding model named CmdCaliper, enabling the computation of semantic similarity with command lines. Performance evaluations demonstrate that the smallest version of Cmd-Caliper (30 million parameters) suppresses state-of-the-art (SOTA) sentence embedding models with ten times more parameters across various tasks (e.g., malicious command-line detection and similar command-line retrieval). Our study explores the feasibility of data generation using LLMs in the cybersecurity domain. Furthermore, we release our proposed command-line dataset, embedding models' weights and all program codes to the public. This advancement paves the way for more effective command-line embedding for future researchers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Improving Text Embeddings with Large Language ModelsLiang Wang, Nan Yang, Xiaolong Huang, Linjun Yang 等ACL 2024
- DISTDET: A Cost-Effective Distributed Cyber Threat Detection SystemFeng Dong, Liu Wang, Xu Nie, Fei Shao 等USENIX Security 2023
相关 Paper
- Toward Cybersecurity-Expert Small Language ModelsMatan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sa 等ICML 2026 · 被引用 7 次
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 被引用 4 次
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesHaoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin 等ACL 2025
- RMCBench: Benchmarking Large Language Models' Resistance to Malicious CodeJiachi Chen, Qingyuan Zhong, Yanlin Wang, Kaiwen Ning 等ASE 2024 · 被引用 6 次
- Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning SystemZiyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang 等ACL 2025
