Automatic Unit Test Generation for Machine Learning Libraries: How Far Are We?
Song Wang, Nishtha Shrestha, Abarna Kucheri Subburaman, Junjie Wang, Moshi Wei, Nachiappan Nagappan
摘要
Automatic unit test generation that explores the input space and produces effective test cases for given programs have been studied for decades. Many unit test generation tools that can help generate unit test cases with high structural coverage over a program have been examined. However, the fact that existing test generation tools are mainly evaluated on general software programs calls into question about its practical effectiveness and usefulness for machine learning libraries, which are statistically-orientated and have fundamentally different nature and construction from general software projects. In this paper, we set out to investigate the effectiveness of existing unit test generation techniques on machine learning libraries. To investigate this issue, we conducted an empirical study on five widely-used machine learning libraries with two popular unit test case generation tools, i.e., EVOSUITE and Randoop. We find that (1) most of the machine learning libraries do not maintain a high-quality unit test suite regarding commonly applied quality metrics such as code coverage (on average is 34.1%) and mutation score (on average is 21.3%), (2) unit test case generation tools, i.e., EVOSUITE and Randoop, lead to clear improvements in code coverage and mutation score, however, the improvement is limited, and (3) there exist common patterns in the uncovered code across the five machine learning libraries that can be used to improve unit test case generation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Bugs in Quantum computing platforms: an empirical studyMatteo Paltenghi, Michael PradelOOPSLA 2022 · 被引用 70 次
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault LocalizationSungmin Kang, Gabin An, Shin YooFSE 2024 · 被引用 69 次
- CLEAR: Contrastive Learning for API RecommendationMoshi Wei, Nima Shiri Harzevili, Yuchao Huang, Junjie Wang 等ICSE 2022 · 被引用 43 次
- Understanding performance problems in deep learning systemsJunming Cao, Bihuan Chen, Chao Sun, Longjie Hu 等FSE 2022 · 被引用 33 次
- Virtual Reality (VR) Automated Testing in the Wild: A Case Study on Unity-Based VR ApplicationsDhia Elhaq Rzig, Nafees Iqbal, Isabella Attisano, Xue Qin 等ISSTA 2023 · 被引用 20 次
相关 Paper
- Effective Unit Test Generation for Java Null Pointer ExceptionsMyungho Lee, Jiseong Bak, Seokhyeon Moon, Yoonchan Jhi 等ASE 2024 · 被引用 1 次
- Increasing the Effectiveness of Automatically Generated Tests by Improving Class ObservabilityGeraldine Galindo-Gutierrez, Juan Pablo Sandoval Alcocer, Nicolas Jimenez-Fuentes, Alexandre Bergel 等ICSE 2025 · 被引用 1 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
- Do Automatic Test Generation Tools Generate Flaky Tests?Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck 等ICSE 2024 · 被引用 12 次
- Domain Adaptation for Code Model-Based Unit Test Case GenerationJiho Shin, Sepehr Hashtroudi, Hadi Hemmati, Song WangISSTA 2024 · 被引用 21 次
