FLEX: fixing flaky tests in machine learning projects by updating assertion bounds
Saikat Dutta, August Shi, Sasa Misailovic
摘要
Many machine learning (ML) algorithms are inherently randommultiple executions using the same inputs may produce slightly different results each time. Randomness impacts how developers write tests that check for end-to-end quality of their implementations of these ML algorithms. In particular, selecting the proper thresholds for comparing obtained quality metrics with the reference results is a non-intuitive task, which may lead to flaky test executions.
We present FLEX, the first tool for automatically fixing flaky tests due to algorithmic randomness in ML algorithms. FLEX fixes tests that use approximate assertions to compare actual and expected values that represent the quality of the outputs of ML algorithms. We present a technique for systematically identifying the acceptable bound between the actual and expected output quality that also minimizes flakiness. Our technique is based on the Peak Over Threshold method from statistical Extreme Value Theory, which estimates the tail distribution of the output values observed from several runs. Based on the tail distribution, FLEX updates the bound used in the test, or selects the number of test re-runs, based on a desired confidence level.
We evaluate FLEX on a corpus of 35 tests collected from the latest versions of 21 ML projects. Overall, FLEX identifies and proposes a fix for 28 tests. We sent 19 pull requests, each fixing one test, to the developers. So far, 9 have been accepted by the developers.
• Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 被引用 91 次
- Repairing Order-Dependent Flaky Tests via Test GenerationChengpeng Li, Chenguang Zhu, Wenxi Wang, August ShiICSE 2022 · 被引用 22 次
- Preempting Flaky Tests via Non-Idempotent-Outcome TestsAnjiang Wei, Pu Yi, Zhengxi Li, Tao Xie 等ICSE 2022 · 被引用 20 次
- TERA: optimizing stochastic regression tests in machine learning projectsSaikat Dutta, Jeeva Selvam, Aryaman Jain, Sasa MisailovicISSTA 2021 · 被引用 12 次
- Understanding and Improving Flaky Test ClassificationShanto Rahman, Saikat Dutta, August ShiOOPSLA 2025 · 被引用 5 次
它引用的顶会 Paper8
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 被引用 107 次
- Audee: Automated Testing for Deep Learning FrameworksQianyu Guo, Xiaofei Xie, Yi Li, Xiaoyu Zhang 等ASE 2020 · 被引用 83 次
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang 等ISSTA 2020 · 被引用 71 次
- Detecting numerical bugs in neural network architecturesYuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong 等FSE 2020 · 被引用 66 次
- Domain-Specific Fixes for Flaky Tests with Wrong Assumptions on Underdetermined SpecificationsPeilun Zhang, Yanjie Jiang, Anjiang Wei, Victoria Stodden 等ICSE 2021 · 被引用 26 次
相关 Paper
- Balancing Effectiveness and Flakiness of Non-Deterministic Machine Learning TestsChunqiu Steven Xia, Saikat Dutta, Sasa Misailovic, Darko Marinov 等ICSE 2023 · 被引用 5 次
- Test Flimsiness: Characterizing Flakiness Induced by Mutation to the Code Under TestOwain Parry, Gregory M. Kapfhammer, Michael Hilton, Phil McMinnICSE 2026
- Neurosymbolic Repair of Test FlakinessYang Chen, Reyhaneh JabbarvandISSTA 2024 · 被引用 8 次
- Higher income, larger loan? monotonicity testing of machine learning modelsArnab Sharma, Heike WehrheimISSTA 2020 · 被引用 12 次
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck 等ICSE 2024 · 被引用 12 次
