TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
Kush Jain, Gabriel Synnaeve, Baptiste Rozière
摘要
Code generation models can help improve many common software tasks ranging from code completion to defect prediction. Most of the existing benchmarks for code generation LLMs focus on code authoring or code completion. Surprisingly, there has been far less effort dedicated to benchmarking software testing, despite the strong correlation between well-tested software and effective bug detection. To address this gap, we create and release TESTGENEVAL, a large-scale benchmark to measure test generation performance. Based on SWEBench, TEST-GENEVAL comprises 68,647 tests from 1,210 code and test file pairs across 11 well-maintained Python repositories. It covers initial tests authoring, test suite completion, and code coverage improvements. Test authoring simulates the process of a developer writing a test suite from scratch, while test completion mimics the scenario where a developer aims to improve the coverage of an existing test suite. We evaluate several popular models, with sizes ranging from 7B to 405B parameters. Our detailed analysis highlights TESTGENEVAL's contribution to a comprehensive evaluation of test generation performance. In particular, models struggle to generate high-coverage test suites, with the best model, GPT-4o, achieving an average coverage of only 35.2%. This is primarily due to models struggling to reason about execution, and their frequent assertion errors when addressing complex code paths.
We provide all the code for our benchmark at https://figshare.com/s/ 51171ae97cd21d233d4f, including detailed instructions on how to run our benchmark, and even extend it. We also provide a website with all model generations for TESTGENEVAL. We hope that this will enable the community to use TESTGENEVAL and further build upon our work.
Our contribution are as follows:
• We release a benchmark for partial and full test suite generation on a realistic set of 1,210 snippets in 11 repositories. We use coverage and mutation score metrics to evaluate the value of the generated test suites
• We evaluate various prominent open and closed-source code generation models on our benchmark, and show that, for large scale repositories, models struggle to generate high coverage test suites
• We release docker images allowing to easily run code from these 11 repositories and evaluate scores on our benchmark
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux 等NeurIPS 2025 · 被引用 291 次
- Rethinking Verification for LLM Code Generation: From Generation to TestingZihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu 等NeurIPS 2025 · 被引用 19 次
- Learning to Generate Unit Test via Adversarial Reinforcement LearningDongjun Lee, Changho Hwang, Kimin LeeICLR 2026 · 被引用 14 次
- How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix PerspectiveXianzhen Luo, Jinyang Huang, Wenzhen Zheng, Qingfu Zhu 等ICLR 2026 · 被引用 5 次
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test GenerationSteven Liu, Jane Luo, Xin Zhang, Aofan Liu 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama 等ICML 2024 · 被引用 270 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
相关 Paper
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 被引用 9 次
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun 等AAAI 2026 · 被引用 10 次
- UTBoost: Rigorous Evaluation of Coding Agents on SWE-BenchBoxi Yu, Yuxuan Zhu, Pinjia He, Daniel KangACL 2025 · 被引用 20 次
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng 等ICLR 2026 · 被引用 27 次
