BlueCodeAgent: A Blue Teaming Agent Powered by Automated Red Teaming for CodeGen AI
Chengquan Guo, Yuzhou Nie, Chulin Xie, Zinan Lin, Wenbo Guo, Bo Li
Abstract
Existing research on CodeGen AI security mainly focuses on red teaming, which aims to uncover vulnerabilities and risks in AI-generated code. However, progress on the blue teaming side remains limited, as effective defenses require a deep security analysis of given tasks and edge cases. To fill in this gap, we propose BlueCodeAgent, an end-to-end blue teaming agent powered by automated red teaming. Our red teaming component generates diverse risky instances, providing effective edge cases and guidance for the subsequent blue teaming process. Our blue teaming agent then conducts multi-level defense, leveraging these red teaming examples to detect previously seen and unseen risk scenarios through constitution summarization and dynamic code analysis. Our evaluation across four representative code-related tasks-bias instruction detection, malicious instruction detection, vulnerable code detection, and prompt injection detection-shows that BlueCodeAgent achieves significant gains over diverse baselines. In particular, for vulnerability detection tasks, BlueCodeAgent integrates dynamic analysis to effectively reduce false positives, a challenging problem as base models tend to be over-conservative. Overall, with GPT-4o as the base model, BlueCodeAgent achieves an average F1 score improvement of 14.7% across four tasks compared to directly prompting the model, attributed to its ability to summarize actionable constitutions and perform dynamic analysis. Our code and data are publicly available at https: //github.com/1mocat/BlueCodeAgent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt et al.S&P 2022 · 725 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce et al.S&P 2024 · 167 citations
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 98 citations
Related papers
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code AgentsChengquan Guo, Chulin Xie, Yu Yang, Zhaorun Chen et al.ICLR 2026 · 13 citations
- RedCoder: Automated Multi-Turn Red Teaming for Code LLMsWenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung et al.ACL 2026 · 6 citations
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM AgentsPengfei He, Ash Fox, Lesly Miculicich, Stefan Friedli et al.ICML 2026 · 10 citations
- AdvAgent: Controllable Blackbox Red-teaming on Web AgentsChejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao et al.ICML 2025
- TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy ExplorationChunxiao Li, Lijun Li, Jing ShaoCVPR 2026 · 4 citations
