Test Flimsiness: Characterizing Flakiness Induced by Mutation to the Code Under Test
Owain Parry, Gregory M. Kapfhammer, Michael Hilton, Phil McMinn
Abstract
Flaky tests, which fail non-deterministically against the same version of code, pose a well-established challenge to software developers. In this paper, we characterize the overlooked phenomenon of test FLIMsiness: FLakiness Induced by Mutations to the code under test. These mutations are generated by the same operators found in standard mutation testing tools. Flimsiness has profound implications for software testing researchers. Previous studies quantified the impact of pre-existing flaky tests on mutation testing, but we reveal that mutations themselves can induce flakiness, exposing a previously neglected threat. This has serious effects beyond mutation testing, calling into question the reliability of any technique that relies on deterministic test outcomes in response to mutations.
On the other hand, flimsiness presents an opportunity to surface potential flakiness that may otherwise remain hidden. Prior work perturbed the execution environment to augment rerunning-based detection and the test code to support benchmarking. We advance these efforts by perturbing a third major source of flakiness: the code under test. We conducted an empirical study on over half a million test suite executions across 28 Python projects. Our statistical analysis on over 30 million mutant-test pairs unveiled flimsiness in 54% of projects. We found that extending the standard rerunning flaky test detection strategy with code-under-test mutations detects a substantially larger number of flaky tests (median 740 vs. 163) and uncovers many that the standard strategy is unlikely to detect.
• Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efcce0c3-b20d-4a27-bd12-bdeec9359ffbBuilds on5
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 63 citations
- FlakiMe: Laboratory-Controlled Test Flakiness Impact AssessmentMaxime Cordy, Renaud Rwemalika, Adriano Franci, Mike Papadakis et al.ICSE 2022 · 14 citations
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck et al.ICSE 2024 · 12 citations
- Does mutation testing improve testing practices?Goran Petrovic, Marko Ivankovic, Gordon Fraser, René JustICSE 2021 · 6 citations
- Transforming Test Suites into CroissantsYang Chen, Alperen Yildiz, Darko Marinov, Reyhaneh JabbarvandISSTA 2023 · 5 citations
Related papers
- A large-scale longitudinal study of flaky testsWing Lam, Stefan Winter, Anjiang Wei, Tao Xie et al.OOPSLA 2020 · 63 citations
- An Empirical Analysis of UI-based Flaky TestsAlan Romano, Zihe Song, Sampath Grandhi, Wei Yang et al.ICSE 2021 · 43 citations
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang et al.ISSTA 2020 · 71 citations
- Ripples of a Mutation - An Empirical Study of Propagation Effects in Mutation TestingHang Du, Vijay Krishna Palepu, James A. JonesICSE 2024 · 6 citations
- To Kill a Mutant: An Empirical Study of Mutation Testing KillsHang Du, Vijay Krishna Palepu, James A. JonesISSTA 2023 · 5 citations
