Change And Cover: Last-Mile, Pull Request-Based Regression Test Augmentation
Zitong Zhou, Matteo Paltenghi, Miryung Kim, Michael Pradel
Abstract
Software is in constant evolution, with developers frequently submitting pull requests (PRs) to introduce new features or fix bugs. Testing newly added or modified code in PRs is critical to maintaining software quality. Yet, even in projects with extensive test suites, some of the code modified in PRs may remain untested, leaving a “last-mile” regression test gap. Existing automated test generators mostly focus on improving overall code coverage, but do not specifically target the uncovered lines in PRs. This paper presents Change And Cover (ChaCo), a novel, LLM-based test augmentation technique that specifically addresses the last-mile regression test gap in PRs. Our approach is enabled by three key contributions: (i) Instead of focusing on overall code coverage, ChaCo considers a specific PR and the lines left uncovered after applying the PR, offering developers augmented tests for code just when it is on the developers’ mind. (ii) We identify providing suitable test context as a crucial challenge for an LLM to generate useful tests, and present two techniques to extract relevant test content, such as existing test functions, fixtures, and data generators. (iii) To make augmented tests acceptable for developers, ChaCo carefully integrates them into the existing test suite, e.g., by matching the test’s structure and style with the existing tests, and generates a summary of the test addition for developer review. We evaluate ChaCo on 145 PRs from three popular, complex, and well-tested open-source projects—SciPy, Qiskit, and Pandas. The approach successfully helps 30% of PRs achieve full patch coverage, at the affordable cost of $0.11 per PR, demonstrating its effectiveness and feasibility. A qualitative assessment of the generated tests shows that human reviewers find the tests to be worth adding (4.53/5.0), well integrated (4.20/5.0), and relevant to the PR (4.70/5.0). Ablation studies show test context is crucial for context-aware test generation, leading to 2 × coverage. In a contribution study, we submitted 12 tests to these projects, of which 8 have already been merged, and two previously unknown bugs were discovered and fixed. We envision our approach to be integrated into CI workflows, automating the last mile of regression test augmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22e9074c-cb14-4ed5-903d-d57fbf6be4b6Cited by top-tier papers1
Ask how each one uses itBuilds on14
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPTChunqiu Steven Xia, Lingming ZhangISSTA 2024 · 105 citations
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 96 citations
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- Regression Greybox FuzzingXiaogang Zhu, Marcel BöhmeCCS 2021 · 84 citations
Related papers
- CoverUp: Effective High Coverage Test Generation for PythonJuan Altmayer Pizzorno, Emery D. BergerFSE 2025 · 13 citations
- Testora: Using Natural Language Intent to Detect Behavioral RegressionsMichael PradelICSE 2026 · 1 citation
- Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based TestingKonstantinos Kitsios, Marco Castelluccio, Alberto BacchelliASE 2025 · 1 citation
- LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test GenerationGwihwan Go, Quan Zhang, Chijin Zhou, Zhao Wei et al.ICSE 2026 · 2 citations
- TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language ModelsCuong Chi Le, Cuong Duc Van, Tung Duy Vu, Minh Vu Thai Pham et al.ICSE 2026 · 1 citation
