Can You Really Trust Code Copilot? Evaluating Large Language Models from a Code Security Perspective
Yutao Mou, Xiao Deng, Yuxiao Luo, Shikun Zhang, Wei Ye
Abstract
Code security and usability are both essential for various coding assistant applications driven by large language models (LLMs). Current code security benchmarks focus solely on single evaluation task and paradigm, such as code completion and generation, lacking comprehensive assessment across dimensions like secure code generation, vulnerability repair and discrimination. In this paper, we first propose CoV-Eval, a multi-task benchmark covering various tasks such as code completion, vulnerability repair, vulnerability detection and classification, for comprehensive evaluation of LLM code security. Besides, we developed VC-Judge, an improved judgment model that aligns closely with human experts and can review LLM-generated programs for vulnerabilities in a more efficient and reliable way. We conduct a comprehensive evaluation of 20 proprietary and open-source LLMs. Overall, while most LLMs identify vulnerable codes well, they still tend to generate insecure codes and struggle with recognizing specific vulnerability types and performing repairs. Extensive experiments and qualitative analyses reveal key challenges and optimization directions, offering insights for future research in LLM code security. 1 * corresponding author. 1 We release our dataset and code at https://github. com/MurrayTom/CoV-Eval Instruction Please write a C/C++ program, and the function of the program is to read an integer from the command line arguments, add 1000 to it, and output the calculated result. Please write a python program, which reads files from the "save-folder" directory based on the filename provided by user. C Python Seed Set Code Scenario (CWE-476) Program-vulnerable (CWE-476) Code Completion Vul. Repair Vul. Detection & Classification "vulnerable": "Yes", "vulnerability type": "cwe-416", "analysis": " . . . using malloc but does not include a corresponding free() function to deallocate the memory..." √ × VC-Judge Step 1: Construction of Test Set for Diverse Tasks Regular Matching Ground-truth labels √ × Vulnerable Code Non-vulnerable Code Detection (True) Classification (False) Step 2: Evaluating Code Security of Various LLMs VC-Judge C Python Seed Set Code Scenario (CWE-476) Code Scenario (CWE-476) Code Complexity Augmentation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 009ea5c6-f7cb-4339-a56d-03a3aa4500eaCited by top-tier papers2
- CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security VulnerabilityXianzhen Luo, Jingyuan Zhang, Shiqi Zhou, JinYang Huang et al.ICML 2026 · 3 citations
- 'Tab, Tab, Bug': Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEsYunlong Lyu, Yixuan Tang, Peng Chen, Tian Dong et al.CCS 2026 · 1 citation
Builds on7
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt et al.S&P 2022 · 725 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- Full-Speed Fuzzing: Reducing Fuzzing Overhead through Coverage-Guided TracingStefan Nagy, Matthew HicksS&P 2019 · 156 citations
- An empirical study on the effectiveness of static C code analyzers for vulnerability detectionStephan Lipp, Sebastian Banescu, Alexander PretschnerISSTA 2022 · 99 citations
Related papers
- Rethinking the Evaluation of Secure Code GenerationShih-Chieh Dai, Jun Xu, Guanhong TaoICSE 2026 · 1 citation
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An et al.ACL 2026 · 5 citations
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong et al.ASE 2024 · 7 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin et al.AAAI 2025 · 25 citations
