From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
Yuling Shi, Songsong Wang, Chengcheng Wan, Min Wang, Xiaodong Gu
Abstract
While large language models have made significant strides in code generation, the pass rate of the generated code is bottlenecked on subtle errors, often requiring human intervention to pass tests, especially for complex problems. Existing LLM-based debugging systems treat generated programs as holistic units, failing to address bugs at multiple levels of granularity, from low-level syntax errors to high-level algorithmic flaws. In this paper, we introduce Multi-Granularity Debugger (MGDebugger), a hierarchical code debugger that isolates, identifies, and resolves bugs at various levels of granularity. MGDebugger decomposes problematic code into a hierarchical tree structure of subfunctions, with each level representing a particular granularity of error. During debugging, it analyzes each subfunction and iteratively resolves bugs in a bottom-up manner. To effectively test each subfunction, we propose an LLM-simulated Python executor, which traces code execution and tracks important variable states to pinpoint errors accurately. Extensive experiments with both open-source and commercial LLMs demonstrate that MGDebugger significantly outperforms existing debugging systems, achieving up to 18.9% improvement in accuracy over seed generations in HumanEval and a 97.6% repair success rate in Hu-manEvalFix. Furthermore, MGDebugger effectively generalizes to real-world software defects, fixing 129 bugs in Defects4J v1.2 and 64 bugs in Defects4J v2.0, outperforming state-of-the-art program repair approaches by 13.2% and 8.5% respectively. These results demonstrate MGDebugger's robustness and effectiveness across different scenarios, establishing it as a powerful tool for closing the last mile of code generation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8a19f1e-6911-441c-8982-368a092d6290Cited by top-tier papers12
- FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature ImplementationWei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao et al.ACL 2025 · 40 citations
- Execution Guided Line-by-Line Code GenerationBoaz Lavon, Shahar Katz, Lior WolfNeurIPS 2025 · 13 citations
- PGS: Effective LLM Code Refinement via Property-Oriented and Structurally Minimal FeedbackLehan He, Zeren Chen, Zhe Zhang, Xiang Gao et al.ICML 2026 · 4 citations
- Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory GraphQiuchen Wang, Shihang Wang, Yu Zeng, Qiang Zhang et al.ICML 2026 · 2 citations
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue ResolutionHan Li, Yuling Shi, Shaoxin Lin, Xiaodong Gu et al.ICSE 2026 · 2 citations
Builds on39
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
Related papers
- UniDebugger: Hierarchical Multi-Agent Framework for Unified Software DebuggingCheryl Lee, Chunqiu Steven Xia, Longji Yang, Jen-tse Huang et al.EMNLP 2025
- NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code DebuggingWeiming Zhang, Qingyao Li, Xinyi Dai, Jizheng Chen et al.EMNLP 2025 · 1 citation
- InspectCoder: Dynamic Analysis-Driven Self Repair through Interactive LLM-Debugger CollaborationYunkun Wang, Yue Zhang, Guochang Li, Chen Zhi et al.OOPSLA 2026 · 1 citation
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li et al.FSE 2024 · 60 citations
- ChatDBG: Augmenting Debugging with Large Language ModelsKyla Levin, Nicolas van Kempen, Emery D. Berger, Stephen N. FreundFSE 2025 · 3 citations
