Does mutation testing improve testing practices?
Goran Petrovic, Marko Ivankovic, Gordon Fraser, René Just
Abstract
Various proxy metrics for test quality have been defined in order to guide developers when writing tests. Code coverage is particularly well established in practice, even though the question of how coverage relates to test quality is a matter of ongoing debate. Mutation testing offers a promising alternative: Artificial defects can identify holes in a test suite, and thus provide concrete suggestions for additional tests. Despite the obvious advantages of mutation testing, it is not yet well established in practice. Until recently, mutation testing tools and techniques simply did not scale to complex systems. Although they now do scale, a remaining obstacle is lack of evidence that writing tests for mutants actually improves test quality. In this paper we aim to fill this gap: By analyzing a large dataset of almost 15 million mutants, we investigate how these mutants influenced developers over time, and how these mutants relate to real faults. Our analyses suggest that developers using mutation testing write more tests, and actively improve their test suites with high quality tests such that fewer mutants remain. By analyzing a dataset of past fixes of real high-priority faults, our analyses further provide evidence that mutants are indeed coupled with real faults. In other words, had mutation testing been used for the changes introducing the faults, it would have reported a live mutant that could have prevented the bug.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Prioritizing Mutants to Guide Mutation TestingSamuel J. Kaufman, Ryan Featherman, Justin Alvin, Bob Kurtz et al.ICSE 2022 · 36 citations
- Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons LearnedSoneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum et al.FSE 2023 · 30 citations
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 12 citations
- Who Judges the Judge: An Empirical Study on Online Judge TestsKaibo Liu, Yudong Han, Jie M. Zhang, Zhenpeng Chen et al.ISSTA 2023 · 10 citations
- Validating SMT Solvers via Skeleton Enumeration Empowered by Historical Bug-Triggering InputsMaolin Sun, Yibiao Yang, Ming Wen, Yongcong Wang et al.ICSE 2023 · 9 citations
Builds on1
Related papers
- On the use of mutation analysis for evaluating student test suite qualityJames Perretta, Andrew DeOrio, Arjun Guha, Jonathan BellISSTA 2022 · 6 citations
- Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)Junda Zhao, Shurui Zhou, Eldan CohenISSTA 2026
- How Does Killing Surviving Mutants Help Detect Real Bugs with Assertion Generation? A Controlled ExperimentHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2026
- State Field Coverage: A Metric for Oracle QualityFacundo Molina, Nazareno Aguirre, Alessandra GorlaASE 2025 · 1 citation
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 5 citations
