CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability Detection
Jujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu, Yulin Yao, Jiaqi Gao, Minjiao Yang
摘要
Software vulnerabilities have increased sharply, underscoring the growing urgency for effective detection methods. Although large language model (LLM) based methods have shown promise in this task, current state-of-the-art LLM approaches struggle with functions that have long contexts. In this paper, we propose CTX-Coder, a context-enhanced vulnerability detection framework that enables LLMs to selectively focus on relevant contextual functions. To achieve this, we represent the contextual functions as embeddings and integrate them with the target code via cross-attention, thereby enhancing the model's ability to capture contextual information. Furthermore, to equip the model with the ability to recognize these embedding features, we propose a two-stage pretraining pipeline. We also introduce a new dataset, CTX-VUL, which addresses the limitations of existing datasets that either lack contextual information for vulnerable functions or are not publicly available. Extensive experiments demonstrate that CTX-Coder (10B) significantly outperforms baseline models with even larger parameters, such as Qwen2.5-14B and SecGPT. As the input code length increases, CTX-Coder’s F1 score drops by only 5.01%, while other models degrade by 25% to 41.5%, showing strong robustness to long-context scenarios and the effectiveness of our design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce 等S&P 2024 · 被引用 167 次
- Self-Supervised Bug Detection and RepairMiltiadis Allamanis, Henry Jackson-Flux, Marc BrockschmidtNeurIPS 2021 · 被引用 145 次
- MVD: Memory-Related Vulnerability Detection Based on Flow-Sensitive Graph Neural NetworksSicong Cao, Xiaobing Sun, Lili Bo, Rongxin Wu 等ICSE 2022 · 被引用 100 次
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin 等ICSE 2025 · 被引用 44 次
相关 Paper
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao 等ISSTA 2025 · 被引用 2 次
- LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language ModelsAhmed Lekssays, Hamza Mouhcine, Khang Tran, Ting Yu 等USENIX Security 2025
- Bridge and Hint: Extending Pre-trained Language Models for Long-Range CodeYujia Chen, Cuiyun Gao, Zezhou Yang, Hongyu Zhang 等ISSTA 2024 · 被引用 4 次
- SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionXin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao 等ISSTA 2024 · 被引用 17 次
- Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?Yu Zhao, Lina Gong, Zhiqiu Huang, Yongwei Wang 等ASE 2024 · 被引用 10 次
