Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code Candidates
Zhihong Sun, Yao Wan, Jia Li, Hongyu Zhang, Zhi Jin, Ge Li, Chen Lyu
摘要
Large Language Models (LLMs), such as GPT-4, StarCoder, and Code Llama, are transforming the way developers approach programming by automatically generating code based on given contexts, such as natural language descriptions or incomplete surrounding code. Despite advancements, generating syntactically and semantically correct code remains challenging, especially for complex programming tasks. Existing approaches typically generate multiple candidate solutions using LLMs to increase the likelihood of producing correct code. However, selecting the correct code from these candidates --- a process known as code ranking --- remains a major challenge. Current research on code ranking can be categorized into execution-based and non-execution-based methods. Execution-based methods, although effective, encounter notable limitations, such as scarcity of quality unit tests and security risks. Non-execution-based methods like CodeRanker, which rely solely on classification labels to train a code ranker, struggle to capture subtle errors and provide detailed error insights. Recognizing the strengths and limitations of both approaches, we propose a new method that integrates the advantages of execution-based and non-execution-based techniques. The key insight of our work is that an effective code ranker is expected to truly comprehend the underlying causes of erroneous code, as relying solely on classification labels is insufficient. Inspired by this, this paper puts forward RankEF, an innovative approach for code ranking that leverages execution feedback. RankEF employs multi-task learning to integrate code classification with execution feedback generation. This approach enables the model to understand the reasons behind incorrect code, distinguishing between correct and incorrect solutions without the need to execute the code during the ranking phase. Experiments on three code generation benchmarks---APPS, MBPP, and HumanEval---demonstrate that RankEF significantly outperforms the state-of-the-art CodeRanker, achieving relative improvements of +30.97%, +31.43%, and +19.51% in Pass@1, Pass@2, and Pass@5 on APPS test, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi 等WWW 2025 · 被引用 38 次
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage DesignsYi Gui, Yao Wan, Zhen Li, Zhongyi Zhang 等WWW 2025 · 被引用 24 次
- Dynamic-Static Synergistic Selection Method for Candidate Code Solutions with Generated Test CasesRenbiao Liu, Jiang-Tian Xue, Chao-Zeng Ma, Hui Sun 等AAAI 2026 · 被引用 2 次
- SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated CodeQinglin Wang, Zhihong Sun, Ruyun Wang, Tao Huang 等ASE 2025 · 被引用 1 次
- LaTCoder: Converting Webpage Design to Code with Layout-as-ThoughtYi Gui, Zhen Li, Zhongyi Zhang, Guohao Wang 等KDD 2025 · 被引用 1 次
它引用的顶会 Paper17
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov 等ICML 2023 · 被引用 318 次
相关 Paper
- Fault-Aware Neural Code RankersJeevana Priya Inala, Chenglong Wang, Mei Yang, Andrés Codas 等NeurIPS 2022 · 被引用 61 次
- Improving Code Generation via Small Language Model-as-a-judgeGiuseppe Crupi, Rosalia Tufano, Gabriele BavotaICSE 2026 · 被引用 1 次
- Execution Guided Line-by-Line Code GenerationBoaz Lavon, Shahar Katz, Lior WolfNeurIPS 2025 · 被引用 13 次
- MapCoder: Multi-Agent Code Generation for Competitive Problem SolvingMd. Ashraful Islam, Mohammed Eunus Ali, Md. Rizwan ParvezACL 2024 · 被引用 29 次
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella 等ICML 2025
