Improving Code Generation via Small Language Model-as-a-judge
Giuseppe Crupi, Rosalia Tufano, Gabriele Bavota
摘要
Large language models (LLMs) have shown remarkable capabilities in automated code generation. While effective for mainstream languages, they may underperform on less common or domainspecific languages, prompting companies to develop in-house code generators. While open-source models can be trained for this, only LLMs with tens of billions of parameters match the performance of commercial tools, demanding costly training and deployment. Recent work proposed supporting code generation with smaller models (SLMs) by generating multiple candidate solutions and using another SLM to select the most likely correct one. The most recent work in this area is the one by Sun et al. [29] presenting RankEF, a T5 model trained to rank code solutions using both execution-based and non-execution-based information. However, Sun et al. do not assess the T5 ranker's classification accuracy, that is, how often it misjudges correct implementations as incorrect or vice versa, leaving open questions about the reliability of LMs as code correctness judges for other tasks (e.g., automated code review). Moreover, their experiments involve relatively old models, making it unclear the extent to which such a methodology would still help companies in cheaply training their own code generators with performance comparable to those of massive LLMs. We present a study addressing these limitations. We train several state-of-the-art SLMs as code correctness judges and assess their ability to discriminate between correct and wrong implementations. We show that modern SLMs outperform RankEF, even without exploiting execution-based information. When used as code rankers, they achieve higher performance gains than RankEF and perform competitively with LLMs 5-25× larger, at a fraction of the cost.
• Software and its engineering → Software notations and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
- Coder Reviewer Reranking for Code GenerationTianyi Zhang, Tao Yu, Tatsunori Hashimoto, Mike Lewis 等ICML 2023 · 被引用 125 次
- CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained ModelsHao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang 等ICSE 2024 · 被引用 107 次
- Fault-Aware Neural Code RankersJeevana Priya Inala, Chenglong Wang, Mei Yang, Andrés Codas 等NeurIPS 2022 · 被引用 61 次
相关 Paper
- Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code CandidatesZhihong Sun, Yao Wan, Jia Li, Hongyu Zhang 等ASE 2024 · 被引用 5 次
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng 等ICLR 2026 · 被引用 27 次
- On LLMs’ Internal Representation of Code CorrectnessFrancisco Ribeiro, Claudio Spiess, Premkumar Devanbu, Sarah NadiICSE 2026
- CodeJudge: Evaluating Code Generation with Large Language ModelsWeixi Tong, Tianyi ZhangEMNLP 2024 · 被引用 25 次
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov 等ICML 2023 · 被引用 318 次
