FLEX: fixing flaky tests in machine learning projects by updating assertion bounds
Saikat Dutta, August Shi, Sasa Misailovic
Abstract
Many machine learning (ML) algorithms are inherently randommultiple executions using the same inputs may produce slightly different results each time. Randomness impacts how developers write tests that check for end-to-end quality of their implementations of these ML algorithms. In particular, selecting the proper thresholds for comparing obtained quality metrics with the reference results is a non-intuitive task, which may lead to flaky test executions.
We present FLEX, the first tool for automatically fixing flaky tests due to algorithmic randomness in ML algorithms. FLEX fixes tests that use approximate assertions to compare actual and expected values that represent the quality of the outputs of ML algorithms. We present a technique for systematically identifying the acceptable bound between the actual and expected output quality that also minimizes flakiness. Our technique is based on the Peak Over Threshold method from statistical Extreme Value Theory, which estimates the tail distribution of the output values observed from several runs. Based on the tail distribution, FLEX updates the bound used in the test, or selects the number of test re-runs, based on a desired confidence level.
We evaluate FLEX on a corpus of 35 tests collected from the latest versions of 21 ML projects. Overall, FLEX identifies and proposes a fix for 28 tests. We sent 19 pull requests, each fixing one test, to the developers. So far, 9 have been accepted by the developers.
• Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28bcac53-cae0-4bab-b039-5aa55a5b2851Cited by top-tier papers6
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 91 citations
- Repairing Order-Dependent Flaky Tests via Test GenerationChengpeng Li, Chenguang Zhu, Wenxi Wang, August ShiICSE 2022 · 22 citations
- Preempting Flaky Tests via Non-Idempotent-Outcome TestsAnjiang Wei, Pu Yi, Zhengxi Li, Tao Xie et al.ICSE 2022 · 20 citations
- TERA: optimizing stochastic regression tests in machine learning projectsSaikat Dutta, Jeeva Selvam, Aryaman Jain, Sasa MisailovicISSTA 2021 · 12 citations
- Understanding and Improving Flaky Test ClassificationShanto Rahman, Saikat Dutta, August ShiOOPSLA 2025 · 5 citations
Builds on8
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 107 citations
- Audee: Automated Testing for Deep Learning FrameworksQianyu Guo, Xiaofei Xie, Yi Li, Xiaoyu Zhang et al.ASE 2020 · 83 citations
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang et al.ISSTA 2020 · 71 citations
- Detecting numerical bugs in neural network architecturesYuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong et al.FSE 2020 · 66 citations
- Domain-Specific Fixes for Flaky Tests with Wrong Assumptions on Underdetermined SpecificationsPeilun Zhang, Yanjie Jiang, Anjiang Wei, Victoria Stodden et al.ICSE 2021 · 26 citations
Related papers
- Balancing Effectiveness and Flakiness of Non-Deterministic Machine Learning TestsChunqiu Steven Xia, Saikat Dutta, Sasa Misailovic, Darko Marinov et al.ICSE 2023 · 5 citations
- Test Flimsiness: Characterizing Flakiness Induced by Mutation to the Code Under TestOwain Parry, Gregory M. Kapfhammer, Michael Hilton, Phil McMinnICSE 2026
- Neurosymbolic Repair of Test FlakinessYang Chen, Reyhaneh JabbarvandISSTA 2024 · 8 citations
- Higher income, larger loan? monotonicity testing of machine learning modelsArnab Sharma, Heike WehrheimISSTA 2020 · 12 citations
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck et al.ICSE 2024 · 12 citations
