Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code Candidates
Zhihong Sun, Yao Wan, Jia Li, Hongyu Zhang, Zhi Jin, Ge Li, Chen Lyu
Abstract
Large Language Models (LLMs), such as GPT-4, StarCoder, and Code Llama, are transforming the way developers approach programming by automatically generating code based on given contexts, such as natural language descriptions or incomplete surrounding code. Despite advancements, generating syntactically and semantically correct code remains challenging, especially for complex programming tasks. Existing approaches typically generate multiple candidate solutions using LLMs to increase the likelihood of producing correct code. However, selecting the correct code from these candidates --- a process known as code ranking --- remains a major challenge. Current research on code ranking can be categorized into execution-based and non-execution-based methods. Execution-based methods, although effective, encounter notable limitations, such as scarcity of quality unit tests and security risks. Non-execution-based methods like CodeRanker, which rely solely on classification labels to train a code ranker, struggle to capture subtle errors and provide detailed error insights. Recognizing the strengths and limitations of both approaches, we propose a new method that integrates the advantages of execution-based and non-execution-based techniques. The key insight of our work is that an effective code ranker is expected to truly comprehend the underlying causes of erroneous code, as relying solely on classification labels is insufficient. Inspired by this, this paper puts forward RankEF, an innovative approach for code ranking that leverages execution feedback. RankEF employs multi-task learning to integrate code classification with execution feedback generation. This approach enables the model to understand the reasons behind incorrect code, distinguishing between correct and incorrect solutions without the need to execute the code during the ranking phase. Experiments on three code generation benchmarks---APPS, MBPP, and HumanEval---demonstrate that RankEF significantly outperforms the state-of-the-art CodeRanker, achieving relative improvements of +30.97%, +31.43%, and +19.51% in Pass@1, Pass@2, and Pass@5 on APPS test, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d15380f-4b81-4ed7-9d7f-be35d42cc67eCited by top-tier papers8
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage DesignsYi Gui, Yao Wan, Zhen Li, Zhongyi Zhang et al.WWW 2025 · 24 citations
- Dynamic-Static Synergistic Selection Method for Candidate Code Solutions with Generated Test CasesRenbiao Liu, Jiang-Tian Xue, Chao-Zeng Ma, Hui Sun et al.AAAI 2026 · 2 citations
- SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated CodeQinglin Wang, Zhihong Sun, Ruyun Wang, Tao Huang et al.ASE 2025 · 1 citation
- LaTCoder: Converting Webpage Design to Code with Layout-as-ThoughtYi Gui, Zhen Li, Zhongyi Zhang, Guohao Wang et al.KDD 2025 · 1 citation
Builds on17
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov et al.ICML 2023 · 318 citations
Related papers
- Fault-Aware Neural Code RankersJeevana Priya Inala, Chenglong Wang, Mei Yang, Andrés Codas et al.NeurIPS 2022 · 61 citations
- Improving Code Generation via Small Language Model-as-a-judgeGiuseppe Crupi, Rosalia Tufano, Gabriele BavotaICSE 2026 · 1 citation
- Execution Guided Line-by-Line Code GenerationBoaz Lavon, Shahar Katz, Lior WolfNeurIPS 2025 · 13 citations
- MapCoder: Multi-Agent Code Generation for Competitive Problem SolvingMd. Ashraful Islam, Mohammed Eunus Ali, Md. Rizwan ParvezACL 2024 · 29 citations
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
