VFCionX: Bridging Large and Small Models for Robust Vulnerability-Fixing Commit Identification
Xing Cui, Jingzheng Wu, Wenxiang Ou, Tianyue Luo, Zhiyuan Li, Xiang Ling
Abstract
Vulnerability-Fixing Commit Identification(VFCI) is a critical task in software security maintenance that aims to automatically identify code commits that patch security vulnerabilities. However, existing approaches face challenges in handling low-quality commit messages and entangled commits, which limit their identification performance. To address these issues, we propose VFCionX, a novel VFCI framework that integrates large and small language models in a collaborative architecture. VFCionX consists of three core modules: Message Classifier, Patch Classifier, and Ensemble Classifier. The Message Classifier employs a multi-source contextual augmentation strategy to enhance the quality of commit messages and fine-tunes the Qwen2.5-1.5B model, significantly improving classification performance in the textual modality. The Patch Classifier combines heuristic rules with a Qwen2.5-Coder-7B-driven file selector to filter noise from entangled commits, and incorporates a line-level feature extractor based on CodeBERT and CNN to capture local pattern differences between added and deleted code lines. The Ensemble Classifier integrates predictions from both channels using the AdaBoost algorithm, enhancing model robustness and generalization. Experimental results on five popular C/C++ repositories comprising 24,630 commits show that VFCionX achieves an F1-score of 81.47%, outperforming the best baseline by 9.42%. Ablation studies validate the effectiveness of each component, while sensitivity analysis reveals optimal parameter settings for balancing performance and noise resilience. This work provides a new and effective solution for robust vulnerability patch identification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4258b30-e2f8-4952-95bf-3ec029442a27Builds on8
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- A Large-Scale Empirical Study of Security PatchesFrank Li, Vern PaxsonCCS 2017 · 273 citations
- Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language ModelsGen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen et al.NeurIPS 2023 · 157 citations
- Deep just-in-time defect prediction: how far are we?Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming ZhangISSTA 2021 · 97 citations
- Demystifying the Vulnerability Propagation and Its Evolution via Dependency Trees in the NPM EcosystemChengwei Liu, Sen Chen, Lingling Fan, Bihuan Chen et al.ICSE 2022 · 94 citations
Related papers
- CLNX: Bridging Code and Natural Language for C/C++ Vulnerability-Contributing Commits IdentificationZeqing Qin, Yiwei Wu, Lansheng HanAAAI 2025 · 2 citations
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLMXu Yang, Wenhan Zhu, Michael Pacheco, Jiayuan Zhou et al.FSE 2025 · 5 citations
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis PipelineZheng Zhang, Haonan Li, Xingyu Li, Hang Zhang et al.NDSS 2026 · 1 citation
- CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability DetectionJujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu et al.AAAI 2026
- Not Every Patch is an Island: LLM-Enhanced Identification of Multiple Vulnerability PatchesYi Song, Dongchen Xie, Lin Xu, He Zhang et al.ASE 2025
