FlakeFlagger: Predicting Flakiness Without Rerunning Tests
Abdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan Bell
Abstract
When developers make changes to their code, they typically run regression tests to detect if their recent changes (re) introduce any bugs. However, many tests are flaky, and their outcomes can change non-deterministically, failing without apparent cause. Flaky tests are a significant nuisance in the development process, since they make it more difficult for developers to trust the outcome of their tests, and hence, it is important to know which tests are flaky. The traditional approach to identify flaky tests is to rerun them multiple times: if a test is observed both passing and failing on the same code, it is definitely flaky. We conducted a very large empirical study looking for flaky tests by rerunning the test suites of 24 projects 10,000 times each, and found that even with this many reruns, some previously identified flaky tests were still not detected. We propose FlakeFlagger, a novel approach that collects a set of features describing the behavior of each test, and then predicts tests that are likely to be flaky based on similar behavioral features. We found that FlakeFlagger correctly labeled as flaky at least as many tests as a state-of-the-art flaky test classifier, but that FlakeFlagger reported far fewer false positives. This lower false positive rate translates directly to saved time for researchers and developers who use the classification result to guide more expensive flaky test detection processes. Evaluated on our dataset of 23 projects with flaky tests, FlakeFlagger outperformed the prior approach (by F1 score) on 16 projects and tied on 4 projects. Our results indicate that this approach can be effective for identifying likely flaky tests prior to running time-consuming flaky test detectors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e626305-da22-446e-a0cb-b4edadb48a71Cited by top-tier papers11
- Preempting Flaky Tests via Non-Idempotent-Outcome TestsAnjiang Wei, Pu Yi, Zhengxi Li, Tao Xie et al.ICSE 2022 · 20 citations
- FlakiMe: Laboratory-Controlled Test Flakiness Impact AssessmentMaxime Cordy, Renaud Rwemalika, Adriano Franci, Mike Papadakis et al.ICSE 2022 · 14 citations
- Evolution-aware detection of order-dependent flaky testsChengpeng Li, August ShiISSTA 2022 · 12 citations
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck et al.ICSE 2024 · 12 citations
- WEFix: Intelligent Automatic Generation of Explicit Waits for Efficient Web End-to-End Flaky TestsXinyue Liu, Zihe Song, Weike Fang, Wei Yang et al.WWW 2024 · 9 citations
Builds on2
Related papers
- Detecting Flaky Tests by Controlling Nondeterministic API BehaviorHengchen Yuan, Jiefang Lin, August ShiOOPSLA 2026
- Dependent-test-aware regression testing techniquesWing Lam, August Shi, Reed Oei, Sai Zhang et al.ISSTA 2020 · 44 citations
- A Dataset of Reproducible Flaky-Test FailuresSuzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan et al.ISSTA 2026
- An Empirical Analysis of UI-based Flaky TestsAlan Romano, Zihe Song, Sampath Grandhi, Wei Yang et al.ICSE 2021 · 43 citations
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 107 citations
