A large-scale longitudinal study of flaky tests
Wing Lam, Stefan Winter, Anjiang Wei, Tao Xie, Darko Marinov, Jonathan Bell
摘要
Flaky tests are tests that can non-deterministically pass or fail for the same code version. These tests undermine regression testing efficiency, because developers cannot easily identify whether a test fails due to their recent changes or due to flakiness. Ideally, one would detect flaky tests right when flakiness is introduced, so that developers can then immediately remove the flakiness. Some software organizations, e.g., Mozilla and Netflix, run some toolsÐdetectorsÐto detect flaky tests as soon as possible. However, detecting flaky tests is costly due to their inherent non-determinism, so even state-of-the-art detectors are often impractical to be used on all tests for each project change. To combat the high cost of applying detectors, these organizations typically run a detector solely on newly added or directly modified tests, i.e., not on unmodified tests or when other changes occur (including changes to the test suite, the code under test, and library dependencies). However, it is unclear how many flaky tests can be detected or missed by applying detectors in only these limited circumstances.
To better understand this problem, we conduct a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky. We apply two state-of-theart detectors to 55 Java projects, identifying a total of 245 flaky tests that can be compiled and run in the code version where each test was added. We find that 75% of flaky tests (184 out of 245) are flaky when added, indicating substantial potential value for developers to run detectors specifically on newly added tests. However, running detectors solely on newly added tests would still miss detecting 25% of flaky tests. The percentage of flaky tests that can be detected does increase to 85% when detectors are run on newly added or directly modified tests. The remaining 15% of flaky tests become flaky due to other changes and can be detected only when detectors are always applied to all tests. Our study is the first to empirically evaluate when tests become flaky and to recommend guidelines for applying detectors in the future. CCS Concepts: • Software and its engineering → Software testing and debugging; Software evolution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 被引用 91 次
- FlakeFlagger: Predicting Flakiness Without Rerunning TestsAbdulrahman Alshammari, Christopher Morris, Michael Hilton, Jonathan BellICSE 2021 · 被引用 63 次
- Domain-Specific Fixes for Flaky Tests with Wrong Assumptions on Underdetermined SpecificationsPeilun Zhang, Yanjie Jiang, Anjiang Wei, Victoria Stodden 等ICSE 2021 · 被引用 26 次
- Repairing Order-Dependent Flaky Tests via Test GenerationChengpeng Li, Chenguang Zhu, Wenxi Wang, August ShiICSE 2022 · 被引用 22 次
它引用的顶会 Paper1
相关 Paper
- Evolution-aware detection of order-dependent flaky testsChengpeng Li, August ShiISSTA 2022 · 被引用 12 次
- Preempting Flaky Tests via Non-Idempotent-Outcome TestsAnjiang Wei, Pu Yi, Zhengxi Li, Tao Xie 等ICSE 2022 · 被引用 20 次
- An Empirical Analysis of UI-based Flaky TestsAlan Romano, Zihe Song, Sampath Grandhi, Wei Yang 等ICSE 2021 · 被引用 43 次
- A Dataset of Reproducible Flaky-Test FailuresSuzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan 等ISSTA 2026
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 被引用 107 次
