Do Automatic Test Generation Tools Generate Flaky Tests?
Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck, Phil McMinn, Gordon Fraser
摘要
Non-deterministic test behavior, or flakiness, is common and dreaded among developers. Researchers have studied the issue and proposed approaches to mitigate it. However, the vast majority of previous work has only considered developer-written tests. The prevalence and nature of flaky tests produced by test generation tools remain largely unknown. We ask whether such tools also produce flaky tests and how these differ from developer-written ones. Furthermore, we evaluate mechanisms that suppress flaky test generation. We sample 6 356 projects written in Java or Python. For each project, we generate tests using EvoSuite (Java) and Pynguin (Python), and execute each test 200 times, looking for inconsistent outcomes. Our results show that flakiness is at least as common in generated tests as in developer-written tests. Nevertheless, existing flakiness suppression mechanisms implemented in EvoSuite are effective in alleviating this issue (71.7 % fewer flaky tests). Compared to developer-written flaky tests, the causes of generated flaky tests are distributed differently. Their non-deterministic behavior is more frequently caused by randomness, rather than by networking and concurrency. Using flakiness suppression, the remaining flaky tests differ significantly from any flakiness previously reported, where most are attributable to runtime optimizations and EvoSuiteinternal resource thresholds. These insights, with the accompanying dataset, can help maintainers to improve test generation tools, give recommendations for developers using these tools, and serve as a foundation for future research in test flakiness or test generation. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CoverUp: Effective High Coverage Test Generation for PythonJuan Altmayer Pizzorno, Emery D. BergerFSE 2025 · 被引用 13 次
- Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based TestingKonstantinos Kitsios, Marco Castelluccio, Alberto BacchelliASE 2025 · 被引用 1 次
- Test Flimsiness: Characterizing Flakiness Induced by Mutation to the Code Under TestOwain Parry, Gregory M. Kapfhammer, Michael Hilton, Phil McMinnICSE 2026
它引用的顶会 Paper7
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 被引用 107 次
- Detecting flaky tests in probabilistic and machine learning applicationsSaikat Dutta, August Shi, Rutvik Choudhary, Zhekun Zhang 等ISSTA 2020 · 被引用 71 次
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 被引用 63 次
- Empirically revisiting and enhancing IR-based test-case prioritizationQianyang Peng, August Shi, Lingming ZhangISSTA 2020 · 被引用 49 次
- Dependent-test-aware regression testing techniquesWing Lam, August Shi, Reed Oei, Sai Zhang 等ISSTA 2020 · 被引用 44 次
相关 Paper
- A Dataset of Reproducible Flaky-Test FailuresSuzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan 等ISSTA 2026
- Automatic Unit Test Generation for Machine Learning Libraries: How Far Are We?Song Wang, Nishtha Shrestha, Abarna Kucheri Subburaman, Junjie Wang 等ICSE 2021 · 被引用 36 次
- An Empirical Analysis of UI-based Flaky TestsAlan Romano, Zihe Song, Sampath Grandhi, Wei Yang 等ICSE 2021 · 被引用 43 次
- Transforming Test Suites into CroissantsYang Chen, Alperen Yildiz, Darko Marinov, Reyhaneh JabbarvandISSTA 2023 · 被引用 5 次
- FlakiMe: Laboratory-Controlled Test Flakiness Impact AssessmentMaxime Cordy, Renaud Rwemalika, Adriano Franci, Mike Papadakis 等ICSE 2022 · 被引用 14 次
