Lune

ISSTA2026顶会

LLMutantKiller: Using Large Language Models to Generate Tests That Kill Mutants

Farideh Khalili, Aidan Domondon, Harshit Garg, Frank Tip

2026年份

摘要

The primary goal of mutation testing is to assess the quality of an application’s test suite. This is accomplished by introducing syntactic changes into a program and determining if any test failures occur for the resulting mutated program, commonly referred to as a mutant. If so, the mutant is said to be killed, confirming that the test suite is of sufficient quality to detect the introduced fault. A problem arises if a mutant does not impact the behavior of any test. Such a surviving mutant may occur for two reasons: either it involves a semantics- preserving program transformation or the test suite is not strong enough. Determining why a mutant survives often involves complex, non-local reasoning. This paper presents an LLM-based test generation technique for killing surviving mutants, implemented in a tool called LLMutantKiller. The technique is feedback-directed in the sense that if a test is produced that does not kill a given mutant, the LLM is re-prompted up to a specified number of times with scenario-specific feedback such as syntax errors, dependency violations, or execution logs (e.g., failing assertions) and asked to try again. We evaluate LLMutantKiller on 915 randomly selected surviving mutants produced by StrykerJS, a state-of-the-art mutation testing tool, across 13 open-source JavaScript/TypeScript applications. The results show that LLMutantKiller kills up to 95.3% of the surviving mutants classified as inducing behavioral changes and that it rarely produces invalid tests.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖