Test Flimsiness: Characterizing Flakiness Induced by Mutation to the Code Under Test
Owain Parry, Gregory M. Kapfhammer, Michael Hilton, Phil McMinn
摘要
Flaky tests, which fail non-deterministically against the same version of code, pose a well-established challenge to software developers. In this paper, we characterize the overlooked phenomenon of test FLIMsiness: FLakiness Induced by Mutations to the code under test. These mutations are generated by the same operators found in standard mutation testing tools. Flimsiness has profound implications for software testing researchers. Previous studies quantified the impact of pre-existing flaky tests on mutation testing, but we reveal that mutations themselves can induce flakiness, exposing a previously neglected threat. This has serious effects beyond mutation testing, calling into question the reliability of any technique that relies on deterministic test outcomes in response to mutations.
On the other hand, flimsiness presents an opportunity to surface potential flakiness that may otherwise remain hidden. Prior work perturbed the execution environment to augment rerunning-based detection and the test code to support benchmarking. We advance these efforts by perturbing a third major source of flakiness: the code under test. We conducted an empirical study on over half a million test suite executions across 28 Python projects. Our statistical analysis on over 30 million mutant-test pairs unveiled flimsiness in 54% of projects. We found that extending the standard rerunning flaky test detection strategy with code-under-test mutations detects a substantially larger number of flaky tests (median 740 vs. 163) and uncovers many that the standard strategy is unlikely to detect.
• Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 被引用 63 次
- FlakiMe: Laboratory-Controlled Test Flakiness Impact AssessmentMaxime Cordy, Renaud Rwemalika, Adriano Franci, Mike Papadakis 等ICSE 2022 · 被引用 14 次
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck 等ICSE 2024 · 被引用 12 次
- Does mutation testing improve testing practices?Goran Petrovic, Marko Ivankovic, Gordon Fraser, René JustICSE 2021 · 被引用 6 次
- Transforming Test Suites into CroissantsYang Chen, Alperen Yildiz, Darko Marinov, Reyhaneh JabbarvandISSTA 2023 · 被引用 5 次
相关 Paper
- A large-scale longitudinal study of flaky testsWing Lam, Stefan Winter, Anjiang Wei, Tao Xie 等OOPSLA 2020 · 被引用 63 次
- An Empirical Analysis of UI-based Flaky TestsAlan Romano, Zihe Song, Sampath Grandhi, Wei Yang 等ICSE 2021 · 被引用 43 次
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang 等ISSTA 2020 · 被引用 71 次
- Ripples of a Mutation - An Empirical Study of Propagation Effects in Mutation TestingHang Du, Vijay Krishna Palepu, James A. JonesICSE 2024 · 被引用 6 次
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 被引用 5 次
