On the Reproducibility of Software Defect Datasets
Hao-Nan Zhu, Cindy Rubio-González
摘要
Software defect datasets are crucial to facilitating the evaluation and comparison of techniques in fields such as fault localization, test generation, and automated program repair. However, the reproducibility of software defect artifacts is not immune to breakage. In this paper, we conduct a study on the reproducibility of software defect artifacts. First, we study five state-of-the-art Java defect datasets. Despite the multiple strategies applied by dataset maintainers to ensure reproducibility, all datasets are prone to breakages. Second, we conduct a case study in which we systematically test the reproducibility of 1,795 software artifacts during a 13-month period. We find that 62.6% of the artifacts break at least once, and 15.3% artifacts break multiple times. We manually investigate the root causes of breakages and handcraft 10 patches, which are automatically applied to 1,055 distinct artifacts in 2,948 fixes. Based on the nature of the root causes, we propose automated dependency caching and artifact isolation to prevent further breakage. In particular, we show that isolating artifacts to eliminate external dependencies increases reproducibility to 95% or higher, which is on par with the level of reproducibility exhibited by the most reliable manually curated dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding Industry Perspectives of Static Application Security Testing (SAST) EvaluationYuan Li, Peisen Yao, Kan Yu, Chengpeng Wang 等FSE 2025 · 被引用 1 次
- Defects4REST: A Benchmark of Real-World Defects to Enable Controlled Testing and Debugging Studies for REST APIsRahil Mehta, Pushpak Katkhede, Manish MotwaniICSE 2026 · 被引用 1 次
- BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMsDi Wu, Xu He, Shu Wang, Kun SunUSENIX Security 2026
它引用的顶会 Paper6
- Fixing dependency errors for Python build reproducibilitySuchita Mukherjee, Abigail Almanza, Cindy Rubio-GonzálezISSTA 2021 · 被引用 55 次
- Extracting Concise Bug-Fixing Patches from Human-Written Patches in Version Control SystemsYanjie Jiang, Hui Liu, Nan Niu, Lu Zhang 等ICSE 2021 · 被引用 38 次
- Shipwright: A Human-in-the-Loop System for Dockerfile RepairJordan Henkel, Denini Silva, Leopoldo Teixeira, Marcelo d'Amorim 等ICSE 2021 · 被引用 31 次
- Detecting and reproducing error-code propagation bugs in MPI implementationsDaniel DeFreez, Antara Bhowmick, Ignacio Laguna, Cindy Rubio-GonzálezPPoPP 2020 · 被引用 13 次
- Testing self-adaptive software with probabilistic guarantees on performance metricsClaudio Mandrioli, Martina MaggioFSE 2020 · 被引用 12 次
相关 Paper
- RegMiner: towards constructing a large regression dataset from code evolution historyXuezhi Song, Yun Lin, Siang Hwee Ng, Yijian Wu 等ISSTA 2022 · 被引用 11 次
- On the efficiency of test suite based program repair: A Systematic Assessment of 16 Automated Repair Systems for Java ProgramsKui Liu, Shangwen Wang, Anil Koyuncu, Kisub Kim 等ICSE 2020 · 被引用 116 次
- A Dataset of Reproducible Flaky-Test FailuresSuzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan 等ISSTA 2026
- A Large-Scale Empirical Review of Patch Correctness Checking ApproachesJun Yang, Yuehan Wang, Yiling Lou, Ming Wen 等FSE 2023 · 被引用 11 次
- SoK: Automated Vulnerability Repair: Methods, Tools, and AssessmentsYiwei Hu, Zhen Li, Kedie Shu, Shenghua Guan 等USENIX Security 2025
