Contextual Predictive Mutation Testing
Kush Jain, Uri Alon, Alex Groce, Claire Le Goues
Abstract
Mutation testing is a powerful technique for assessing and improving test suite quality that artificially introduces bugs and checks whether the test suites catch them. However, it is also computationally expensive and thus does not scale to large systems and projects. One promising recent approach to tackling this scalability problem uses machine learning to predict whether the tests will detect the synthetic bugs, without actually running those tests. However, existing predictive mutation testing approaches still misclassify 33% of detection outcomes on a randomly sampled set of mutant-test suite pairs. We introduce MutationBERT, an approach for predictive mutation testing that simultaneously encodes the source method mutation and test method, capturing key context in the input representation. Thanks to its higher precision, Mu-tationBERT saves 33% of the time spent by a prior approach on checking/verifying live mutants. MutationBERT, also outperforms the state-of-the-art in both same project and cross project settings, with meaningful improvements in precision, recall, and F1 score. We validate our input representation, and aggregation approaches for lifting predictions from the test matrix level to the test suite level, finding similar improvements in performance. MutationBERT not only enhances the state-of-the-art in predictive mutation testing, but also presents practical benefits for real-world applications, both in saving developer time and finding hard to detect mutants.
• Software and its engineering → Dynamic analysis; Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4535767c-85b6-493e-beb6-57da5ece7981Cited by top-tier papers1
Ask how each one uses itBuilds on4
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
- Prioritizing Mutants to Guide Mutation TestingSamuel J. Kaufman, Ryan Featherman, Justin Alvin, Bob Kurtz et al.ICSE 2022 · 36 citations
- Program merge conflict resolution via neural transformersAlexey Svyatkovskiy, Sarah Fakhoury, Negar Ghorbani, Todd Mytkowicz et al.FSE 2022 · 33 citations
Related papers
- Using Active Learning to Train Predictive Mutation Testing with Minimal DataMiklos BorsiASE 2025
- Re-evaluating Detection of Equivalent Mutants using LLMs: We Should Properly Measure How Far We AreArjun Tandon, Mehmet Fırat Dündar, Milkiyas Gebremichael Gebru, Darko Marinov et al.ISSTA 2026
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program RepairHaoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu et al.ASE 2020 · 81 citations
- MuDelta: Delta-Oriented Mutation Testing at Commit TimeWei Ma, Thierry Titcheu Chekam, Mike Papadakis, Mark HarmanICSE 2021 · 10 citations
- Automatic Unit Test Generation for Machine Learning Libraries: How Far Are We?Song Wang, Nishtha Shrestha, Abarna Kucheri Subburaman, Junjie Wang et al.ICSE 2021 · 36 citations
