Fault-Aware Neural Code Rankers
Jeevana Priya Inala, Chenglong Wang, Mei Yang, Andrés Codas, Mark Encarnación, Shuvendu K. Lahiri, Madanlal Musuvathi, Jianfeng Gao
摘要
Large language models (LLMs) have demonstrated an impressive ability to generate code for various programming tasks. In many instances, LLMs can generate a correct program for a task when given numerous trials. Consequently, a recent trend is to do large scale sampling of programs using a model and then filtering/ranking the programs based on the program execution on a small number of known unit tests to select one candidate solution. However, these approaches assume that the unit tests are given and assume the ability to safely execute the generated programs (which can do arbitrary dangerous operations such as file manipulations). Both of the above assumptions are impractical in real-world software development. In this paper, we propose CODERANKER, a neural ranker that can predict the correctness of a sampled program without executing it. Our CODERANKER is fault-aware i.e., it is trained to predict different kinds of execution information such as predicting the exact compile/runtime error type (e.g., an IndexError or a TypeError). We show that CODERANKER can significantly increase the pass@1 accuracy of various code generation models (including Codex [11] , GPT-Neo, GPT-J) on APPS [25] , HumanEval [11] and MBPP [3] datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Is Self-Repair a Silver Bullet for Code Generation?Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao 等ICLR 2024 · 被引用 195 次
- Coder Reviewer Reranking for Code GenerationTianyi Zhang, Tao Yu, Tatsunori Hashimoto, Mike Lewis 等ICML 2023 · 被引用 125 次
- CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modulesHung Le, Hailin Chen, Amrita Saha, Akash Gokul 等ICLR 2024 · 被引用 73 次
- ALGO: Synthesizing Algorithmic Programs with Generated Oracle VerifiersKexun Zhang, Danqing Wang, Jingtao Xia, William Yang Wang 等NeurIPS 2023 · 被引用 68 次
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan 等ICLR 2023 · 被引用 64 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- InCoder: A Generative Model for Code Infilling and SynthesisDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang 等ICLR 2023 · 被引用 140 次
- Latent Execution for Neural Program Synthesis Beyond Domain-Specific LanguagesXinyun Chen, Dawn Song, Yuandong TianNeurIPS 2021 · 被引用 56 次
- Neural Program Generation Modulo Static AnalysisRohan Mukherjee, Yeming Wen, Dipak Chaudhari, Thomas W. Reps 等NeurIPS 2021 · 被引用 28 次
相关 Paper
- Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code CandidatesZhihong Sun, Yao Wan, Jia Li, Hongyu Zhang 等ASE 2024 · 被引用 5 次
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov 等ICML 2023 · 被引用 318 次
- Self-Edit: Fault-Aware Code Editor for Code GenerationKechi Zhang, Zhuo Li, Jia Li, Ge Li 等ACL 2023 · 被引用 42 次
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- Oracle-Guided Program Selection from Large Language ModelsZhiyu Fan, Haifeng Ruan, Sergey Mechtaev, Abhik RoychoudhuryISSTA 2024 · 被引用 4 次
