How Does Killing Surviving Mutants Help Detect Real Bugs with Assertion Generation? A Controlled Experiment
Hang Du, Vijay Krishna Palepu, James A. Jones
Abstract
Killing surviving mutants is a central activity of mutation testing. This activity is motivated by the coupling-effect hypothesis: tests that expose simple artificial faults can also detect more complex, previously unseen real bugs. Despite these claimed benefits, automated studies have not directly measured the causal impact of mutant killing on real-bug detection. This limitation stems from open-ended mutant-killing strategies and a fundamental evaluation asymmetry that obscures causal attribution. In this work, we present the first large-scale controlled experiment that directly measures whether killing surviving mutants, without knowledge of future real bugs, would have enabled their detection. We model mutant killing as a selective, incremental process under realistic budget constraints, and we restrict test improvements to assertion augmentation. This restriction enables precise attribution of each test augmentation to a specific mutant-killing action. To support the experiment, we design a fully automated, fault-based assertion-augmentation technique that operates uniformly on mutants and real bugs and integrate it into Defects4J. Our controlled experiment yields several key empirical insights: (1) Across 642 Defects4J bugs, we find that 104 bugs would become detectable by adding an additional assertion to an existing passing, non-triggering test. (2) When coupling exists, a real bug is, on average, coupled with 21 surviving mutants, through which mutant killing can produce triggering tests. This number is substantially higher than the average of two mutants reported in prior studies. In those studies, coupling is inferred solely from documented bug-fixing tests rather than from tests derived via mutant killing. (3) Among these bugs, 63 of the 104 are detectable through principled mutant-killing (test augmentation) process. Notably, killing a randomly selected 30% of relevant surviving mutants, using only one assertion per mutant, suffices to detect 84.5% of these bugs. (4) By substituting mutants with real bugs and comparing their resulting assertion augmentation outputs, we find that real bugs induce broader behavioral effects than mutants, affecting more memory state locations, variables, and tests. (5) When mutation-derived assertions detect real bugs, they validate program outputs that overlap with, and are often strict subsets of, those affected by the real bugs. This offers a mechanistic explanation for why killing simple mutants can enable the detection of more complex real bugs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d8a108d5-766f-4d8a-ada2-529797cb7e3cRelated papers
- Leveraging Propagated Infection to Crossfire MutantsHang Du, Vijay Krishna Palepu, James A. JonesICSE 2025 · 1 citation
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 5 citations
- Does mutation testing improve testing practices?Goran Petrovic, Marko Ivankovic, Gordon Fraser, René JustICSE 2021 · 6 citations
- LLMutantKiller: Using Large Language Models to Generate Tests That Kill MutantsFarideh Khalili, Aidan Domondon, Harshit Garg, Frank TipISSTA 2026
- Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set SizeYiqun T. Chen, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst et al.ASE 2020 · 48 citations
