CoverUp: Effective High Coverage Test Generation for Python
Juan Altmayer Pizzorno, Emery D. Berger
摘要
Testing is an essential part of software development. Test generation tools attempt to automate the otherwise labor-intensive task of test creation, but generating high-coverage tests remains challenging. This paper proposes CoverUp, a novel approach to driving the generation of high-coverage Python regression tests. CoverUp combines coverage analysis, code context, and feedback in prompts that iteratively guide the LLM to generate tests that improve line and branch coverage. We evaluate our prototype CoverUp implementation across a benchmark of challenging code derived from open-source Python projects and show that CoverUp substantially improves on the state of the art. Compared to CodaMosa, a hybrid search/LLM-based test generator, CoverUp achieves a per-module median line+branch coverage of 80% (vs. 47%). Compared to MuTAP, a mutation-and LLM-based test generator, CoverUp achieves an overall line+branch coverage of 89% (vs. 77%). We also demonstrate that CoverUp's performance stems not only from the LLM used but from the combined effectiveness of its components. CCS Concepts: • Computing methodologies → Artificial intelligence; • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Demystifying LLM-Based Software Engineering AgentsChunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming ZhangFSE 2025 · 被引用 36 次
- Learning to Generate Unit Test via Adversarial Reinforcement LearningDongjun Lee, Changho Hwang, Kimin LeeICLR 2026 · 被引用 14 次
- LLM Test Generation via Iterative Hybrid Program AnalysisSijia Gu, Noor Nashid, Ali MesbahICSE 2026 · 被引用 6 次
- ChangeGuard: Validating Code Changes via Pairwise Learning-Guided ExecutionLars Gröninger, Beatriz Souza, Michael PradelFSE 2025 · 被引用 4 次
- Knowledge Matters: Injecting Project and Testing Knowledge into LLM-based Unit Test GenerationAnji Li, Mingwei Liu, Zhenxi Chen, Zheng Pei 等ICSE 2026 · 被引用 2 次
它引用的顶会 Paper6
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel 等ICSE 2024 · 被引用 155 次
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang 等FSE 2024 · 被引用 68 次
- A large-scale longitudinal study of flaky testsWing Lam, Stefan Winter, Anjiang Wei, Tao Xie 等OOPSLA 2020 · 被引用 63 次
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck 等ICSE 2024 · 被引用 12 次
相关 Paper
- TestTailor: Generating High-Coverage Tests via Path-Proximal Tests with LLMsXiaoxuan Zhou, Yiling Lou, Jinhao Dong, Dan HaoFSE 2026
- MuMuTestUp: Mutation-Based Multi-agent Test Case UpdateDawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang 等ISSTA 2026
- Change And Cover: Last-Mile, Pull Request-Based Regression Test AugmentationZitong Zhou, Matteo Paltenghi, Miryung Kim, Michael PradelICSE 2026
- TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language ModelsCuong Chi Le, Cuong Duc Van, Tung Duy Vu, Minh Vu Thai Pham 等ICSE 2026 · 被引用 1 次
- Beyond Coverage: Automatic Test Suite Augmentation for Enhanced Effectiveness using Large Language ModelsZeyu Lu, Peng Zhang, Yuge Nie, Yibiao Yang 等OOPSLA 2026 · 被引用 1 次
