Equivalent Mutants in the Wild: Identifying and Efficiently Suppressing Equivalent Mutants for Java Programs
Benjamin Kushigian, Samuel J. Kaufman, Ryan Featherman, Hannah Potter, Ardi Madadi, René Just
Abstract
The presence of equivalent mutants has long been considered a major obstacle to the widespread adoption of mutation analysis and mutation testing. This paper presents a study on the types and prevalence of equivalent mutants in real-world Java programs. We conducted a ground-truth analysis of 1,992 mutants, sampled from 7 open source Java projects. Our analysis identi ed 215 equivalent mutants, which we grouped based on two criteria that describe why the mutants are equivalent and how challenging their detection is. From this analysis, we observed that (1) the median equivalent mutant rate across the 7 projects is 2.97%; (2) many equivalent mutants are caused by common programming patterns and their detection is not much more complex than structural pattern matching over an abstract syntax tree. Based on the ndings of our ground-truth analysis, we developed Equivalent Mutant Suppression (EMS), a technique that comprises 10 e cient and targeted analyses. We evaluated EMS on 19 opensource Java projects, comparing the e ectiveness and e ciency of EMS to two variants of Trivial Compiler Equivalence (TCE), the current state of the art in equivalent mutant detection. Additionally, we analyzed all 9,047 equivalent mutants reported by any tool to better understand the types and frequencies of equivalent mutants found. Overall, EMS detects 8,776 equivalent mutants within 325 seconds; TCE detects 2,124 equivalent mutants in 2,938 hours. CCS Concepts • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e3060ef-b46d-410b-a20f-6182013b4793Cited by top-tier papers3
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- Hybrid Fault-Driven Mutation Testing for PythonSaba Alimadadi, Golnaz GharachorluICSE 2026
- MutDafny: A Mutation-Based Approach to Assess Dafny SpecificationsIsabel Amaral, Alexandra Mendes, José CamposICSE 2026
Builds on3
- Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set SizeYiqun T. Chen, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst et al.ASE 2020 · 48 citations
- Prioritizing Mutants to Guide Mutation TestingSamuel J. Kaufman, Ryan Featherman, Justin Alvin, Bob Kurtz et al.ICSE 2022 · 36 citations
- Does mutation testing improve testing practices?Goran Petrovic, Marko Ivankovic, Gordon Fraser, René JustICSE 2021 · 6 citations
Related papers
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao et al.ISSTA 2024 · 12 citations
- On the use of mutation analysis for evaluating student test suite qualityJames Perretta, Andrew DeOrio, Arjun Guha, Jonathan BellISSTA 2022 · 6 citations
- Re-evaluating Detection of Equivalent Mutants using LLMs: We Should Properly Measure How Far We AreArjun Tandon, Mehmet Fırat Dündar, Milkiyas Gebremichael Gebru, Darko Marinov et al.ISSTA 2026
- Cost measures matter for mutation testing study validityGiovani Guizzo, Federica Sarro, Mark HarmanFSE 2020 · 11 citations
- ATM: Black-box Test Case Minimization based on Test Code Similarity and Evolutionary SearchRongqi Pan, Taher Ahmed Ghaleb, Lionel C. BriandICSE 2023 · 22 citations
