Automatic Unit Test Generation for Machine Learning Libraries: How Far Are We?
Song Wang, Nishtha Shrestha, Abarna Kucheri Subburaman, Junjie Wang, Moshi Wei, Nachiappan Nagappan
Abstract
Automatic unit test generation that explores the input space and produces effective test cases for given programs have been studied for decades. Many unit test generation tools that can help generate unit test cases with high structural coverage over a program have been examined. However, the fact that existing test generation tools are mainly evaluated on general software programs calls into question about its practical effectiveness and usefulness for machine learning libraries, which are statistically-orientated and have fundamentally different nature and construction from general software projects. In this paper, we set out to investigate the effectiveness of existing unit test generation techniques on machine learning libraries. To investigate this issue, we conducted an empirical study on five widely-used machine learning libraries with two popular unit test case generation tools, i.e., EVOSUITE and Randoop. We find that (1) most of the machine learning libraries do not maintain a high-quality unit test suite regarding commonly applied quality metrics such as code coverage (on average is 34.1%) and mutation score (on average is 21.3%), (2) unit test case generation tools, i.e., EVOSUITE and Randoop, lead to clear improvements in code coverage and mutation score, however, the improvement is limited, and (3) there exist common patterns in the uncovered code across the five machine learning libraries that can be used to improve unit test case generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8b93882-187a-4b15-b511-aa83bb49ce9aCited by top-tier papers11
- Bugs in Quantum computing platforms: an empirical studyMatteo Paltenghi, Michael PradelOOPSLA 2022 · 70 citations
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault LocalizationSungmin Kang, Gabin An, Shin YooFSE 2024 · 69 citations
- CLEAR: Contrastive Learning for API RecommendationMoshi Wei, Nima Shiri Harzevili, Yuchao Huang, Junjie Wang et al.ICSE 2022 · 43 citations
- Understanding performance problems in deep learning systemsJunming Cao, Bihuan Chen, Chao Sun, Longjie Hu et al.FSE 2022 · 33 citations
- Virtual Reality (VR) Automated Testing in the Wild: A Case Study on Unity-Based VR ApplicationsDhia Elhaq Rzig, Nafees Iqbal, Isabella Attisano, Xue Qin et al.ISSTA 2023 · 20 citations
Related papers
- Effective Unit Test Generation for Java Null Pointer ExceptionsMyungho Lee, Jiseong Bak, Seokhyeon Moon, Yoonchan Jhi et al.ASE 2024 · 1 citation
- Increasing the Effectiveness of Automatically Generated Tests by Improving Class ObservabilityGeraldine Galindo-Gutierrez, Juan Pablo Sandoval Alcocer, Nicolas Jimenez-Fuentes, Alexandre Bergel et al.ICSE 2025 · 1 citation
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck et al.ICSE 2024 · 12 citations
- Domain Adaptation for Code Model-Based Unit Test Case GenerationJiho Shin, Sepehr Hashtroudi, Hadi Hemmati, Song WangISSTA 2024 · 21 citations
