Rethinking Verification for LLM Code Generation: From Generation to Testing
Zihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu, Wenwei Zhang, Minnan Luo, Songyang Zhang, Kai Chen
摘要
Large language models (LLMs) have recently achieved notable success in code-generation benchmarks such as HumanEval and LiveCodeBench. However, a detailed examination reveals that these evaluation suites often comprise only a limited number of homogeneous test cases, resulting in subtle faults going undetected. This not only artificially inflates measured performance but also compromises accurate reward estimation in reinforcement learning frameworks utilizing verifiable rewards (RLVR). To address these critical shortcomings, we systematically investigate the test-case generation (TCG) task by proposing multi-dimensional metrics designed to rigorously quantify test-suite thoroughness. Furthermore, we introduce a human-LLM collaborative method (SAGA), leveraging human programming expertise with LLM reasoning capability, aimed at significantly enhancing both the coverage and the quality of generated test cases. In addition, we develop a TCGBench to facilitate the study of the TCG task. Experiments show that SAGA achieves a detection rate of 90.62% and a verifier accuracy of 32.58% on TCG-Bench. The Verifier Accuracy (Verifier Acc) of the code generation evaluation benchmark synthesized by SAGA is 10.78% higher than that of LiveCodeBench-v6. These results demonstrate the effectiveness of our proposed method. We hope this work contributes to building a scalable foundation for reliable LLM code evaluation, further advancing RLVR in code generation, and paving the way for automated adversarial test synthesis and adaptive benchmark integration. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart UnderstandingMuye Huang, Lingling Zhang, Jie Ma, Han Lai 等NeurIPS 2025 · 被引用 13 次
- AutoCode: LLMs as Problem Setters for Competitive ProgrammingShang Zhou, Zihan Zheng, Kaiyuan Liu, Zeyu Shen 等ICLR 2026 · 被引用 12 次
- CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsJingwei Shi, Xinxiang Yin, Jing Huang, Shengyu Tao 等ACL 2026 · 被引用 6 次
- How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix PerspectiveXianzhen Luo, Jinyang Huang, Wenzhen Zheng, Qingfu Zhu 等ICLR 2026 · 被引用 5 次
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao 等EMNLP 2025
它引用的顶会 Paper11
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
相关 Paper
- Themis: Automated Constraint-Aware Test Synthesis Framework for Code Reinforcement LearningShengyu Ye, Qi Liu, Hao Jiang, Zheng Zhang 等AAAI 2026
- QiMeng-CodeV-R1: Reasoning-Enhanced Verilog GenerationYaoyu Zhu, Di Huang, Han-Qi Lyu, Xiaoyun Zhang 等NeurIPS 2025 · 被引用 46 次
- HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithimic CodingZhongmou He, Yee Man Choi, Kexun Zhang, Ivan Bercovich 等ICLR 2026
- CODERL+: Improving Code Generation via Reinforcement with Execution Semantics AlignmentXue Jiang, Yihong Dong, Mengyang Liu, Hongyi Deng 等ACL 2026 · 被引用 18 次
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li 等ICLR 2026 · 被引用 4 次
