Balancing Effectiveness and Flakiness of Non-Deterministic Machine Learning Tests
Chunqiu Steven Xia, Saikat Dutta, Sasa Misailovic, Darko Marinov, Lingming Zhang
Abstract
Testing Machine Learning (ML) projects is challenging due to inherent non-determinism of various ML algorithms and the lack of reliable ways to compute reference results. Developers typically rely on their intuition when writing tests to check whether ML algorithms produce accurate results. However, this approach leads to conservative choices in selecting assertion bounds for comparing actual and expected results in test assertions. Because developers want to avoid false positive failures in tests, they often set the bounds to be too loose, potentially leading to missing critical bugs. We present FASER - the first systematic approach for balancing the trade-off between the fault-detection effectiveness and flakiness of non-deterministic tests by computing optimal assertion bounds. FASER frames this trade-off as an optimization problem between these competing objectives by varying the assertion bound. FASER leverages 1) statistical methods to estimate the flakiness rate, and 2) mutation testing to estimate the fault-detection effectiveness. We evaluate FASER on 87 non-deterministic tests collected from 22 popular ML projects. FASER finds that 23 out of 87 studied tests have conservative bounds and proposes tighter assertion bounds that maximizes the fault-detection effectiveness of the tests while limiting flakiness. We have sent 19 pull requests to developers, each fixing one test, out of which 14 pull requests have already been accepted.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio et al.ICSE 2020 · 281 citations
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis et al.ICSE 2020 · 111 citations
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 91 citations
- NNSmith: Generating Diverse and Valid Test Cases for Deep Learning CompilersJiawei Liu, Jinkun Lin, Fabian Ruffy, Cheng Tan et al.ASPLOS 2023 · 90 citations
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 84 citations
Related papers
- FLEX: fixing flaky tests in machine learning projects by updating assertion boundsSaikat Dutta, August Shi, Sasa MisailovicFSE 2021 · 33 citations
- TERA: optimizing stochastic regression tests in machine learning projectsSaikat Dutta, Jeeva Selvam, Aryaman Jain, Sasa MisailovicISSTA 2021 · 12 citations
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 107 citations
- Transforming Test Suites into CroissantsYang Chen, Alperen Yildiz, Darko Marinov, Reyhaneh JabbarvandISSTA 2023 · 5 citations
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 63 citations
