Re-evaluating Detection of Equivalent Mutants using LLMs: We Should Properly Measure How Far We Are
Arjun Tandon, Mehmet Fırat Dündar, Milkiyas Gebremichael Gebru, Darko Marinov, Yiling Lou, Wenxi Wang
摘要
Mutation testing is a widely used approach for measuring test-suite quality. A critical problem in mutation testing is equivalent mutant detection (EMD), i.e., determining if a mutant semantically behaves the same as the original code despite some syntactic differences. A recent study has shown that LLM-based EMD techniques hold great promise, reporting substantial improvements over traditional compiler- and machine-learning–based approaches. In this work, we revisit those recent results and evaluate the generalization capabilities of the proposed LLM-based EMD techniques across two additional datasets that differ from the prior dataset in mutation operators, programming languages, or source projects. Contrary to prior findings, the proposed LLM-based EMD techniques suffer substantial performance degradation on the two additional datasets. Through an extensive analysis, we identify a key factor underlying the differences as original-method–level data leakage (i.e., the same original method appearing in both training and testing sets), indicating that prior results under within-method evaluation do not generalize to cross-method evaluation. We find that the studied LLMs tend to rely on a method-wise majority-voting shortcut rather than reasoning about the semantic effects of mutations. Based on these findings, we call for the adoption of realistic cross-method evaluation and the development of mutation-centric semantic reasoning in future LLM-based EMD research.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao 等ISSTA 2024 · 被引用 12 次
- Learning to Construct Better Mutation FaultsZhao Tian, Junjie Chen, Qihao Zhu, Junjie Yang 等ASE 2022 · 被引用 35 次
- Beyond Coverage: Automatic Test Suite Augmentation for Enhanced Effectiveness using Large Language ModelsZeyu Lu, Peng Zhang, Yuge Nie, Yibiao Yang 等OOPSLA 2026 · 被引用 1 次
- Spotting Code Mutation for Predictive Mutation TestingYifan Zhao, Yizhou Chen, Zeyu Sun, Qingyuan Liang 等ASE 2024 · 被引用 2 次
- LLMutantKiller: Using Large Language Models to Generate Tests That Kill MutantsFarideh Khalili, Aidan Domondon, Harshit Garg, Frank TipISSTA 2026
