Unsupervised Labeling and Extraction of Phrase-based Concepts in Vulnerability Descriptions
Sofonias Yitagesu, Zhenchang Xing, Xiaowang Zhang, Zhiyong Feng, Xiaohong Li, Linyi Han
Abstract
People usually describe the key characteristics of software vulnerabilities in natural language mixed with domain-specific names and concepts. This textual nature poses a significant challenge for the automatic analysis of vulnerabilities. Automatic extraction of key vulnerability aspects is highly desirable but demands significant effort to manually label data for model training. In this paper, we propose an unsupervised approach to label and extract important vulnerability concepts in textural vulnerability descriptions (TVDs). We focus on three types of phrase-based vulnerability concepts (root cause, attack vector, and impact) as they are much more difficult to label and extract than name- or number-based entities (i.e., vendor, product, and version). Our approach is based on a key observation that the same-type of phrases, no matter how they differ in sentence structures and phrase expressions, usually share syntactically similar paths in the sentence parsing trees. Therefore, we propose two path representations (absolute paths and relative paths) and use an auto-encoder to encode such syntactic similarities. To address the discrete nature of our paths, we enhance traditional Variational Auto-encoder (VAE) with Gumble-Max trick for categorical data distribution, and thus creates a Categorical VAE (CaVAE). In the latent space of absolute and relative paths, we further use FIt-TSNE and clustering techniques to generate clusters of the same-type of concepts. Our evaluation confirms the effectiveness of our CaVAE for encoding path representations and the accuracy of vulnerability concepts in the resulting clusters. In a concept classification task, our unsupervisedly labeled vulnerability concepts outperform the two manually labeled datasets from previous work.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get db8d7344-8c68-4c74-a7dd-5e34cbf535bbCited by top-tier papers2
- Identifying Affected Libraries and Their Ecosystems for Open Source Software VulnerabilitiesSusheng Wu, Wenyan Song, Kaifeng Huang, Bihuan Chen et al.ICSE 2024 · 9 citations
- Keyword Extraction From Specification Documents for Planning Security MechanismsJeffy Jahfar Poozhithara, Hazeline U. Asuncion, Brent LagesseICSE 2023 · 1 citation
Related papers
- Teaching AI the 'Why' and 'How' of Software Vulnerability FixesAmiao Gao, Zenong Zhang, Simin Wang, Liguo Huang et al.FSE 2025
- An Empirical Study of Fine-Grained Entity Relationships for Tracing Natural Language and Code Vulnerability ArtifactsSimin Wang, LiGuo Huang, Shiyi Wei, Amiao Gao et al.ICSE 2026 · 1 citation
- Towards the Detection of Inconsistencies in Public Security Vulnerability ReportsYing Dong, Wenbo Guo, Yueqi Chen, Xinyu Xing et al.USENIX Security 2019 · 149 citations
- Learning to Locate and Describe VulnerabilitiesJian Zhang, Shangqing Liu, Xu Wang, Tianlin Li et al.ASE 2023 · 8 citations
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi et al.USENIX Security 2024 · 13 citations
