Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
Bo Lin, Shangwen Wang, Yihao Qin, Liqian Chen, Xiaoguang Mao
摘要
Retrieval-Augmented Code Generation (RACG) leverages external knowledge to enhance Large Language Models (LLMs) in code synthesis, improving the functional correctness of the generated code. However, existing RACG systems largely overlook security, leading to substantial risks. Especially, the poisoning of malicious code into knowledge bases can mislead LLMs, resulting in the generation of insecure outputs, which poses a critical threat in modern software development. To address this, we propose a security-hardening framework for RACG systems, CodeGuarder, that shifts the paradigm from retrieving only functional code examples to incorporating both functional code and security knowledge. Our framework constructs a security knowledge base by analyzing real-world vulnerabilities from the ReposVul dataset. For each code generation query, a retriever decomposes the query into fine-grained sub-tasks and fetches relevant security knowledge. To prioritize critical security guidance, we introduce a re-ranking and filtering mechanism by leveraging the LLMs' susceptibility to different vulnerability types. This filtered security knowledge is seamlessly integrated into the generation prompt. Our evaluation shows CodeGuarder significantly improves code security rates across various LLMs, achieving average improvements of 20.12% in standard RACG, and 31.53% and 21.91% under two distinct poisoning scenarios without compromising functional correctness. Furthermore, CodeGuarder demonstrates strong generalization, enhancing security even when the targeted language's security knowledge is lacking. This work presents CodeGuarder as a pivotal advancement towards building secure and trustworthy RACG systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RESCUE: Retrieval Augmented Secure Code GenerationJiahao Shi, Tianyi ZhangICLR 2026 · 被引用 15 次
- Confundo: Learning to Generate Robust Poison for Practical RAG SystemsHaoyang Hu, Zhejun Jiang, Yueming Lyu, Junyuan Zhang 等USENIX Security 2026 · 被引用 5 次
- Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based BenchmarkingJiexin Wang, Liuwen Cao, Xitong Luo, Yang Cao 等ISSTA 2026 · 被引用 4 次
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li 等FSE 2026 · 被引用 1 次
- Securing Retrieval-Augmented Code Generation via Contextual Knowledge Injection: A Case for Embedded IoT ApplicationsTong Sun, Jingyi Su, Yi Gao, Wei DongUSENIX Security 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 被引用 264 次
相关 Paper
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation SystemsHaowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li 等AAAI 2026 · 被引用 3 次
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu 等ISSTA 2024 · 被引用 10 次
- SRACG: A Code Generation Framework with Selective Retrieval AugmentationMengzhen Wang, Shukai Ma, Songwen Gong, Jiexin Wang 等AAAI 2026
- ImportSnare: Directed 'Code Manual' Hijacking in Retrieval-Augmented Code GenerationKai Ye, Liangcai Su, Chenxiong QianCCS 2025
- Understanding and Improving Model Editing for Secure Code GenerationWeifeng Sun, Quanjun Zhang, Yuchen Chen, Chengran Yang 等ISSTA 2026
