Re-evaluating Detection of Equivalent Mutants using LLMs: We Should Properly Measure How Far We Are
Arjun Tandon, Mehmet Fırat Dündar, Milkiyas Gebremichael Gebru, Darko Marinov, Yiling Lou, Wenxi Wang
Abstract
Mutation testing is a widely used approach for measuring test-suite quality. A critical problem in mutation testing is equivalent mutant detection (EMD), i.e., determining if a mutant semantically behaves the same as the original code despite some syntactic differences. A recent study has shown that LLM-based EMD techniques hold great promise, reporting substantial improvements over traditional compiler- and machine-learning–based approaches. In this work, we revisit those recent results and evaluate the generalization capabilities of the proposed LLM-based EMD techniques across two additional datasets that differ from the prior dataset in mutation operators, programming languages, or source projects. Contrary to prior findings, the proposed LLM-based EMD techniques suffer substantial performance degradation on the two additional datasets. Through an extensive analysis, we identify a key factor underlying the differences as original-method–level data leakage (i.e., the same original method appearing in both training and testing sets), indicating that prior results under within-method evaluation do not generalize to cross-method evaluation. We find that the studied LLMs tend to rely on a method-wise majority-voting shortcut rather than reasoning about the semantic effects of mutations. Based on these findings, we call for the adoption of realistic cross-method evaluation and the development of mutation-centric semantic reasoning in future LLM-based EMD research.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao et al.ISSTA 2024 · 12 citations
- Learning to Construct Better Mutation FaultsZhao Tian, Junjie Chen, Qihao Zhu, Junjie Yang et al.ASE 2022 · 35 citations
- Beyond Coverage: Automatic Test Suite Augmentation for Enhanced Effectiveness using Large Language ModelsZeyu Lu, Peng Zhang, Yuge Nie, Yibiao Yang et al.OOPSLA 2026 · 1 citation
- Spotting Code Mutation for Predictive Mutation TestingYifan Zhao, Yizhou Chen, Zeyu Sun, Qingyuan Liang et al.ASE 2024 · 2 citations
- LLMutantKiller: Using Large Language Models to Generate Tests That Kill MutantsFarideh Khalili, Aidan Domondon, Harshit Garg, Frank TipISSTA 2026
