TERA: optimizing stochastic regression tests in machine learning projects
Saikat Dutta, Jeeva Selvam, Aryaman Jain, Sasa Misailovic
Abstract
The stochastic nature of many Machine Learning (ML) algorithms makes testing of ML tools and libraries challenging. ML algorithms allow a developer to control their accuracy and run-time through a set of hyper-parameters, which are typically manually selected in tests. This choice is often too conservative and leads to slow test executions, thereby increasing the cost of regression testing. We propose TERA, the first automated technique for reducing the cost of regression testing in Machine Learning tools and libraries (jointly referred to as projects) without making the tests more flaky. TERA solves the problem of exploring the trade-off space between execution time of the test and its flakiness as an instance of Stochastic Optimization over the space of algorithm hyper-parameters. TERA presents how to leverage statistical convergence-testing techniques to estimate the level of flakiness of the test for a specific choice of hyper-parameters during optimization. We evaluate TERA on a corpus of 160 tests selected from 15 popular machine learning projects. Overall, TERA obtains a geomean speedup of 2.23x over the original tests, for the minimum passing probability threshold of 99%. We also show that the new tests did not reduce fault detection ability through a mutation study and a study on a set of 12 historical build failures in studied projects. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79b26ec-cee7-4663-94e4-231bbc89a088Cited by top-tier papers3
- FLEX: fixing flaky tests in machine learning projects by updating assertion boundsSaikat Dutta, August Shi, Sasa MisailovicFSE 2021 · 33 citations
- Understanding and Improving Flaky Test ClassificationShanto Rahman, Saikat Dutta, August ShiOOPSLA 2025 · 5 citations
- Balancing Effectiveness and Flakiness of Non-Deterministic Machine Learning TestsChunqiu Steven Xia, Saikat Dutta, Sasa Misailovic, Darko Marinov et al.ICSE 2023 · 5 citations
Builds on8
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 107 citations
- Audee: Automated Testing for Deep Learning FrameworksQianyu Guo, Xiaofei Xie, Yi Li, Xiaoyu Zhang et al.ASE 2020 · 83 citations
- Efficient Compiler Autotuning via Bayesian OptimizationJunjie Chen, Ningxin Xu, Peiqi Chen, Hongyu ZhangICSE 2021 · 73 citations
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang et al.ISSTA 2020 · 71 citations
- Detecting numerical bugs in neural network architecturesYuhao Zhang, Luyao Ren, Liqian Chen, Yingfei Xiong et al.FSE 2020 · 66 citations
Related papers
- Dependent-test-aware regression testing techniquesWing Lam, August Shi, Reed Oei, Sai Zhang et al.ISSTA 2020 · 44 citations
- More Precise Regression Test Selection via Reasoning about Semantics-Modifying ChangesYu Liu, Jiyang Zhang, Pengyu Nie, Milos Gligoric et al.ISSTA 2023 · 19 citations
- Fairness-aware Configuration of Machine Learning LibrariesSaeid Tizpaz-Niari, Ashish Kumar, Gang Tan, Ashutosh TrivediICSE 2022 · 44 citations
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 63 citations
- Test Selection for Unified Regression TestingShuai Wang, Xinyu Lian, Darko Marinov, Tianyin XuICSE 2023 · 9 citations
