Early Life Cycle Software Defect Prediction. Why? How?
N. C. Shrikanth, Suvodeep Majumder, Tim Menzies
摘要
Many researchers assume that, for software analytics, "more data is better." We write to show that, at least for learning defect predictors, this may not be true. To demonstrate this, we analyzed hundreds of popular GitHub projects. These projects ran for 84 months and contained 3,728 commits (median values). Across these projects, most of the defects occur very early in their life cycle. Hence, defect predictors learned from the first 150 commits and four months perform just as well as anything else. This means that, at least for the projects studied here, after the first few months, we need not continually update our defect prediction models. We hope these results inspire other researchers to adopt a "simplicity-first" approach to their work. Some domains require a complex and data-hungry analysis. But before assuming complexity, it is prudent to check the raw data looking for "short cuts" that can simplify the analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- An investigation of cross-project learning in online just-in-time software defect predictionSadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral 等ICSE 2020 · 被引用 49 次
- A Practical Human Labeling Method for Online Just-in-Time Software Defect PredictionLiyan Song, Leandro L. Minku, Cong Teng, Xin YaoFSE 2023 · 被引用 6 次
- FRUGAL: Unlocking Semi-Supervised Learning for Software AnalyticsHuy Tu, Tim MenziesASE 2021 · 被引用 9 次
- On the use of evaluation measures for defect prediction studiesRebecca Moussa, Federica SarroISSTA 2022 · 被引用 39 次
- Measuring the Effects of Stack Overflow Code Snippet Evolution on Open-Source Software SecurityAlfusainey Jallow, Michael Schilling, Michael Backes, Sven BugielS&P 2024 · 被引用 6 次
