VFCionX: Bridging Large and Small Models for Robust Vulnerability-Fixing Commit Identification
Xing Cui, Jingzheng Wu, Wenxiang Ou, Tianyue Luo, Zhiyuan Li, Xiang Ling
摘要
Vulnerability-Fixing Commit Identification(VFCI) is a critical task in software security maintenance that aims to automatically identify code commits that patch security vulnerabilities. However, existing approaches face challenges in handling low-quality commit messages and entangled commits, which limit their identification performance. To address these issues, we propose VFCionX, a novel VFCI framework that integrates large and small language models in a collaborative architecture. VFCionX consists of three core modules: Message Classifier, Patch Classifier, and Ensemble Classifier. The Message Classifier employs a multi-source contextual augmentation strategy to enhance the quality of commit messages and fine-tunes the Qwen2.5-1.5B model, significantly improving classification performance in the textual modality. The Patch Classifier combines heuristic rules with a Qwen2.5-Coder-7B-driven file selector to filter noise from entangled commits, and incorporates a line-level feature extractor based on CodeBERT and CNN to capture local pattern differences between added and deleted code lines. The Ensemble Classifier integrates predictions from both channels using the AdaBoost algorithm, enhancing model robustness and generalization. Experimental results on five popular C/C++ repositories comprising 24,630 commits show that VFCionX achieves an F1-score of 81.47%, outperforming the best baseline by 9.42%. Ablation studies validate the effectiveness of each component, while sensitivity analysis reveals optimal parameter settings for balancing performance and noise resilience. This work provides a new and effective solution for robust vulnerability patch identification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- A Large-Scale Empirical Study of Security PatchesFrank Li, Vern PaxsonCCS 2017 · 被引用 273 次
- Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language ModelsGen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen 等NeurIPS 2023 · 被引用 157 次
- Deep just-in-time defect prediction: how far are we?Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming ZhangISSTA 2021 · 被引用 97 次
- Demystifying the Vulnerability Propagation and Its Evolution via Dependency Trees in the NPM EcosystemChengwei Liu, Sen Chen, Lingling Fan, Bihuan Chen 等ICSE 2022 · 被引用 94 次
相关 Paper
- CLNX: Bridging Code and Natural Language for C/C++ Vulnerability-Contributing Commits IdentificationZeqing Qin, Yiwei Wu, Lansheng HanAAAI 2025 · 被引用 2 次
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLMXu Yang, Wenhan Zhu, Michael Pacheco, Jiayuan Zhou 等FSE 2025 · 被引用 5 次
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis PipelineZheng Zhang, Haonan Li, Xingyu Li, Hang Zhang 等NDSS 2026 · 被引用 1 次
- CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability DetectionJujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu 等AAAI 2026
- Not Every Patch is an Island: LLM-Enhanced Identification of Multiple Vulnerability PatchesYi Song, Dongchen Xie, Lin Xu, He Zhang 等ASE 2025
