CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
Yuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao, Qian Chen, Wen Wang, Ziyang Luo, Lei Ma, Dawn Song
Abstract
Large Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requirements. This phenomenon of hallucinations in the code domain has not been systematically explored. To advance the community's understanding and research on this issue, we introduce the concept of code hallucinations and propose a classification method for code hallucination based on execution verification. We categorize code hallucinations into four main types: mapping, naming, resource, and logic hallucinations, with each category further divided into different subcategories to understand and address the unique challenges faced by LLMs in code generation with finer granularity. Additionally, we present a dynamic detection algorithm called Code-Halu designed to detect and quantify code hallucinations. We also introduce the CodeHaluEval benchmark, which includes 8,883 samples from 699 tasks, to systematically and quantitatively evaluate code hallucinations. By evaluating 17 popular LLMs using this benchmark, we reveal significant differences in their accuracy and reliability in code generation, offering detailed insights for further improving the code generation capabilities of LLMs. The CodeHalu benchmark and code are publicly available at https://github.com/yuchen814/CodeHalu .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2067ded9-e12a-4741-b578-11c0539ddb1dCited by top-tier papers6
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel GenerationXinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen et al.AAAI 2026 · 9 citations
- "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking PartnerSzeyi Chan, Jiachen Li, Bingsheng Yao, Amama Mahmood et al.CSCW 2025 · 4 citations
- Intention Chain-of-Thought Prompting with Dynamic Routing for Code GenerationShen Li, Li Huang, Shaoxiong Zhan, Weifeng Sun et al.AAAI 2026 · 1 citation
- HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based FuzzingYukai Zhao, Menghan Wu, Xing Hu, Xin XiaASE 2025 · 1 citation
- ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationSuyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee et al.ACL 2026 · 1 citation
Builds on6
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationJin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai et al.NeurIPS 2022 · 135 citations
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
- Learning Source Phrase Representations for Neural Machine TranslationHongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu et al.ACL 2020 · 18 citations
Related papers
- HALoGEN: Fantastic LLM Hallucinations and Where to Find ThemAbhilasha Ravichander, Shrusti Ghela, David Wadden, Yejin ChoiACL 2025 · 35 citations
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng et al.ACL 2024 · 49 citations
- HalluLens: LLM Hallucination BenchmarkYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn et al.ACL 2025
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi et al.ISSTA 2025 · 53 citations
- Hallucinations in LLM-Based Code Summarization: Unveiling, Detection, and MitigationGuanghua Wan, Yuanning Feng, Yao Wan, Zhaoyang Chu et al.FSE 2026
