Large Language Models of Code Fail at Completing Code with Potential Bugs
Tuan Dinh, Jinman Zhao, Samson Tan, Renato Negrinho, Leonard Lausen, Sheng Zha, George Karypis
Abstract
Large language models of code (Code-LLMs) have recently brought tremendous advances to code completion, a fundamental feature of programming assistance and code intelligence. However, most existing works ignore the possible presence of bugs in the code context for generation, which are inevitable in software development. Therefore, we introduce and study the buggy-code completion problem, inspired by the realistic scenario of real-time code suggestion where the code context contains potential bugs -anti-patterns that can become bugs in the completed program. To systematically study the task, we introduce two datasets: one with synthetic bugs derived from semantics-altering operator changes (buggy-HumanEval) and one with realistic bugs derived from user submissions to coding problems (buggy-FixEval). We find that the presence of potential bugs significantly degrades the generation performance of the high-performing Code-LLMs. For instance, the passing rates of CODEGEN-2B-MONO on test cases of buggy-HumanEval drop more than 50% given a single potential bug in the context. Finally, we investigate several post-hoc methods for mitigating the adverse effect of potential bugs and find that there remains a significant gap in post-mitigation performance. 3 * Equal contribution. † Work done while interning at Amazon Web Services. 3 Code and datasets are available at https://github.com/amazon-science/buggy-code-completion 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Improving Chain-of-Thought Reasoning via Quasi-Symbolic AbstractionsLeonardo Ranaldi, Marco Valentino, André FreitasACL 2025 · 29 citations
- EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code EditsWayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal et al.ICLR 2026 · 7 citations
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language ModelsKunpeng Zhang, Shuai Wang, Jitao Han, Xiaogang Zhu et al.ICSE 2025 · 6 citations
- ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate ModelYang Gu, Jian Cao, Hengyu You, Nengjun Zhu et al.AAAI 2025 · 1 citation
Builds on16
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li et al.ISSTA 2020 · 325 citations
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo et al.ACL 2022 · 208 citations
- DLFix: context-based code transformation learning for automated program repairYi Li, Shaohua Wang, Tien N. NguyenICSE 2020 · 201 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
Related papers
- Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs during Code AdaptationTanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yuxin Zhao et al.FSE 2026
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong et al.ASE 2024 · 7 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Examining Zero-Shot Vulnerability Repair with Large Language ModelsHammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri et al.S&P 2023
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin et al.AAAI 2025 · 25 citations
