A Dataset of Reproducible Flaky-Test Failures
Suzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan, August Shi, Wing Lam
摘要
Flaky tests pass and fail non-deterministically when run on the same version of code. Although many techniques have been proposed to detect, debug, and repair flaky tests, reproducing their failures remains a major challenge due to their inherent nondeterminism. Many datasets related to flaky tests exist to help researchers study them, but these datasets are often composed of disjoint sets of flaky tests, where each dataset provides some unique information over the others, such as flaky tests of many different categories, failure logs of flaky tests, or flaky tests reported by developers vs. flaky tests found by automated tools. In this work, we aim to create a reproducible dataset of flaky tests, which are curated from both developer issue reports and a popular dataset of flaky tests.
Compared to prior flaky test datasets, our dataset is the first to provide (1) a reproducible environment to compile the flaky tests, (2) scripts to run the tests to reproduce the failure, (3) scripts to automatically apply the flaky test fixes and ensure that the test is no longer flaky, and (4) execution logs of the flaky test passing and failing. We present ReproFlake, a dataset of 1115 reproducible flaky tests, spread across four different flaky test categories. We create guidelines to help others contribute to this reproducible dataset, as well as demonstrate how to use our dataset to understand the challenges with reproducing flaky test failures (i.e., challenges that researchers may face when using any of the prior flaky test datasets), the characteristics (e.g., location of the fix and its correlation with the flaky test category), as well as the difficulties researchers may face in using our dataset to collect additional information (e.g., code coverage) about flaky tests. Our findings show that error information helps identify flaky test categories and guide repairs, that unresolved compilation failures highlight challenges in building legacy projects, and that knowing typical fix locations helps prioritize repair efforts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 被引用 107 次
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 被引用 63 次
- Repairing Order-Dependent Flaky Tests via Test GenerationChengpeng Li, Chenguang Zhu, Wenxi Wang, August ShiICSE 2022 · 被引用 22 次
- Preempting Flaky Tests via Non-Idempotent-Outcome TestsAnjiang Wei, Pu Yi, Zhengxi Li, Tao Xie 等ICSE 2022 · 被引用 20 次
- Systematically Producing Test Orders to Detect Order-Dependent Flaky TestsChengpeng Li, Mohammad Mahdi Khosravi, Wing Lam, August ShiISSTA 2023 · 被引用 10 次
相关 Paper
- An Empirical Analysis of UI-based Flaky TestsAlan Romano, Zihe Song, Sampath Grandhi, Wei Yang 等ICSE 2021 · 被引用 43 次
- FlakiMe: Laboratory-Controlled Test Flakiness Impact AssessmentMaxime Cordy, Renaud Rwemalika, Adriano Franci, Mike Papadakis 等ICSE 2022 · 被引用 14 次
- FlakeSync: Automatically Repairing Async Flaky TestsShanto Rahman, August ShiICSE 2024 · 被引用 7 次
- A large-scale longitudinal study of flaky testsWing Lam, Stefan Winter, Anjiang Wei, Tao Xie 等OOPSLA 2020 · 被引用 63 次
- RegMiner: towards constructing a large regression dataset from code evolution historyXuezhi Song, Yun Lin, Siang Hwee Ng, Yijian Wu 等ISSTA 2022 · 被引用 11 次
