How disabled tests manifest in test maintainability challenges?
Dong Jae Kim, Bo Yang, Jinqiu Yang, Tse-Hsun (Peter) Chen
摘要
Software testing is an essential software quality assurance practice. Testing helps expose faults earlier, allowing developers to repair the code and reduce future maintenance costs. However, repairing (i.e., making failing tests pass) may not always be done immediately. Bugs may require multiple rounds of repairs and even remain unfixed due to the difficulty of bug-fixing tasks. To help test maintenance, along with code comments, the majority of testing frameworks (e.g., JUnit and TestNG) have also introduced annotations such as @Ignore to disable failing tests temporarily. Although disabling tests may help alleviate maintenance difficulties, they may also introduce technical debt. With the faster release of applications in modern software development, disabling tests may become the salvation for many developers to meet project deliverables. In the end, disabled tests may become outdated and a source of technical debt, harming long-term maintenance. Despite its harmful implications, there is little empirical research evidence on the prevalence, evolution, and maintenance of disabling tests in practice. To fill this gap, we perform the first empirical study on test disabling practice. We develop a tool to mine 122K commits and detect 3,111 changes that disable tests from 15 open-source Java systems. Our main findings are: (1) Test disabling changes are 19% more common than regular test refactorings, such as renames and type changes.
(2) Our life-cycle analysis shows that 41% of disabled tests are never brought back to evaluate software quality, and most disabled tests stay disabled for several years. (3) We unveil the motivations behind test disabling practice and the associated technical debt by manually studying evolutions of 349 unique disabled tests, achieving a 95% confidence level and a 5% confidence interval. Finally, we present some actionable implications for researchers and developers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A First Look at the Inheritance-Induced Redundant Test ExecutionDong Jae Kim, Jinqiu Yang, Tse-Hsun ChenICSE 2024 · 被引用 2 次
- PairSmell: A Novel Perspective Inspecting Software Modular StructureChenxing Zhong, Daniel Feitosa, Paris Avgeriou, Huang Huang 等ICSE 2025
它引用的顶会 Paper4
- A study on the lifecycle of flaky testsWing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh ThummalapentaICSE 2020 · 被引用 107 次
- CodeShovel: Constructing Method-Level Source Code HistoriesFelix Grund, Shaiful Alam Chowdhury, Nick C. Bradley, Braxton Hall 等ICSE 2021 · 被引用 33 次
- Studying Test Annotation Maintenance in the WildDong Jae Kim, Nikolaos Tsantalis, Tse-Hsun Peter Chen, Jinqiu YangICSE 2021 · 被引用 20 次
- Understanding type changes in JavaAmeya Ketkar, Nikolaos Tsantalis, Danny DigFSE 2020 · 被引用 17 次
相关 Paper
- Paired Code Smells and Test Smells: A Fine-Grained Longitudinal Empirical StudyZiwen CaiISSTA 2026
- 23 shades of self-admitted technical debt: an empirical study on machine learning softwareDavid O'Brien, Sumon Biswas, Sayem Imtiaz, Rabe Abdalkareem 等FSE 2022 · 被引用 37 次
- Evaluating Unit Testing Practices in R PackagesMelina C. VidoniICSE 2021 · 被引用 14 次
- An Empirical Study of Refactorings and Technical Debt in Machine Learning SystemsYiming Tang, Raffi Khatchadourian, Mehdi Bagherzadeh, Rhia Singh 等ICSE 2021 · 被引用 60 次
- Refactorings and Technical Debt in Docker Projects: An Empirical StudyEmna Ksontini, Marouane Kessentini, Thiago do Nascimento Ferreira, Foyzul HassanASE 2021 · 被引用 17 次
