SAGE: Signal-Amplified Guided Embeddings for Vulnerability Detection
Zhengyang Shan, Xu Qian, Jiayun Xin, Minghui Xu, Yue Zhang, Zhen Yang, Hao Wu, Xiuzhen Cheng
Abstract
Software vulnerabilities are a primary threat to modern infrastructure. While static analysis and Graph Neural Networks have long served as the foundation for vulnerability detection, the emergence of Large Language Models (LLMs) has introduced a transformative paradigm driven by superior semantic reasoning and cross-environment generalization. However, in the context of LLM-based vulnerability detection, we identify a fundamental bottleneck in these models termed Signal Submersion : a state where features related to vulnerability are activated internally but numerically overwhelmed by dominant functional semantics. To address this, we propose SAGE ( S ignal- A mplified G uided E mbeddings), a framework that shifts from passive signal submersion to active signal recovery. SAGE integrates task-conditional Sparse Autoencoders (SAEs) to isolate and amplify these faint vulnerability signals. Extensive evaluations on BigVul, PrimeVul, and PreciseBugs demonstrate that SAGE achieves state-of-the-art performance. Notably, SAGE mitigates Signal Submersion by increasing the internal Signal-to-Noise Ratio (SNR) by 12.7× via sparse manifold projection. This mechanistic intervention enables a 7B model to achieve up to 318% Matthews Correlation Coefficient (MCC) gains on unseen distributions and a 319% gain on classic datasets. By maintaining robust performance across 13 programming languages and outperforming 34B baselines, SAGE establishes a more efficient and scalable path to software security than simple parameter scaling.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Towards More Accurate Static Analysis for Taint-Style Bug Detection in Linux KernelHaonan Li, Hang Zhang, Kexin Pei, Zhiyun QianASE 2025 · 5 citations
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin et al.ICSE 2025 · 44 citations
- Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability DetectionBenjamin Steenhoek, Hongyang Gao, Wei LeICSE 2024 · 54 citations
- LOSVER: Line-Level Modifiability Signal-Guided Vulnerability Detection and ClassificationDoha Nam, Jongmoon BaikASE 2025
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao et al.ISSTA 2025 · 2 citations
