LLMutantKiller: Using Large Language Models to Generate Tests That Kill Mutants
Farideh Khalili, Aidan Domondon, Harshit Garg, Frank Tip
Abstract
The primary goal of mutation testing is to assess the quality of an application’s test suite. This is accomplished by introducing syntactic changes into a program and determining if any test failures occur for the resulting mutated program, commonly referred to as a mutant. If so, the mutant is said to be killed, confirming that the test suite is of sufficient quality to detect the introduced fault. A problem arises if a mutant does not impact the behavior of any test. Such a surviving mutant may occur for two reasons: either it involves a semantics- preserving program transformation or the test suite is not strong enough. Determining why a mutant survives often involves complex, non-local reasoning. This paper presents an LLM-based test generation technique for killing surviving mutants, implemented in a tool called LLMutantKiller. The technique is feedback-directed in the sense that if a test is produced that does not kill a given mutant, the LLM is re-prompted up to a specified number of times with scenario-specific feedback such as syntax errors, dependency violations, or execution logs (e.g., failing assertions) and asked to try again. We evaluate LLMutantKiller on 915 randomly selected surviving mutants produced by StrykerJS, a state-of-the-art mutation testing tool, across 13 open-source JavaScript/TypeScript applications. The results show that LLMutantKiller kills up to 95.3% of the surviving mutants classified as inducing behavioral changes and that it rarely produces invalid tests.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b48b1403-9665-48bb-82c6-1ab6f9bfafeeRelated papers
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 5 citations
- Beyond Coverage: Automatic Test Suite Augmentation for Enhanced Effectiveness using Large Language ModelsZeyu Lu, Peng Zhang, Yuge Nie, Yibiao Yang et al.OOPSLA 2026 · 1 citation
- Re-evaluating Detection of Equivalent Mutants using LLMs: We Should Properly Measure How Far We AreArjun Tandon, Mehmet Fırat Dündar, Milkiyas Gebremichael Gebru, Darko Marinov et al.ISSTA 2026
- Leveraging Propagated Infection to Crossfire MutantsHang Du, Vijay Krishna Palepu, James A. JonesICSE 2025 · 1 citation
- Fuzzing JavaScript Engines with Aspect-preserving MutationSoyeon Park, Wen Xu, Insu Yun, Daehee Jang et al.S&P 2020 · 126 citations
