On the Reproducibility of Software Defect Datasets
Hao-Nan Zhu, Cindy Rubio-González
Abstract
Software defect datasets are crucial to facilitating the evaluation and comparison of techniques in fields such as fault localization, test generation, and automated program repair. However, the reproducibility of software defect artifacts is not immune to breakage. In this paper, we conduct a study on the reproducibility of software defect artifacts. First, we study five state-of-the-art Java defect datasets. Despite the multiple strategies applied by dataset maintainers to ensure reproducibility, all datasets are prone to breakages. Second, we conduct a case study in which we systematically test the reproducibility of 1,795 software artifacts during a 13-month period. We find that 62.6% of the artifacts break at least once, and 15.3% artifacts break multiple times. We manually investigate the root causes of breakages and handcraft 10 patches, which are automatically applied to 1,055 distinct artifacts in 2,948 fixes. Based on the nature of the root causes, we propose automated dependency caching and artifact isolation to prevent further breakage. In particular, we show that isolating artifacts to eliminate external dependencies increases reproducibility to 95% or higher, which is on par with the level of reproducibility exhibited by the most reliable manually curated dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17d0f5a7-7f1c-4a8b-a1d6-f3441fbf1d28Cited by top-tier papers3
- Understanding Industry Perspectives of Static Application Security Testing (SAST) EvaluationYuan Li, Peisen Yao, Kan Yu, Chengpeng Wang et al.FSE 2025 · 1 citation
- Defects4REST: A Benchmark of Real-World Defects to Enable Controlled Testing and Debugging Studies for REST APIsRahil Mehta, Pushpak Katkhede, Manish MotwaniICSE 2026 · 1 citation
- BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMsDi Wu, Xu He, Shu Wang, Kun SunUSENIX Security 2026
Builds on6
- Fixing dependency errors for Python build reproducibilitySuchita Mukherjee, Abigail Almanza, Cindy Rubio-GonzálezISSTA 2021 · 55 citations
- Extracting Concise Bug-Fixing Patches from Human-Written Patches in Version Control SystemsYanjie Jiang, Hui Liu, Nan Niu, Lu Zhang et al.ICSE 2021 · 38 citations
- Shipwright: A Human-in-the-Loop System for Dockerfile RepairJordan Henkel, Denini Silva, Leopoldo Teixeira, Marcelo d'Amorim et al.ICSE 2021 · 31 citations
- Detecting and reproducing error-code propagation bugs in MPI implementationsDaniel DeFreez, Antara Bhowmick, Ignacio Laguna, Cindy Rubio-GonzálezPPoPP 2020 · 13 citations
- Testing self-adaptive software with probabilistic guarantees on performance metricsClaudio Mandrioli, Martina MaggioFSE 2020 · 12 citations
Related papers
- RegMiner: towards constructing a large regression dataset from code evolution historyXuezhi Song, Yun Lin, Siang Hwee Ng, Yijian Wu et al.ISSTA 2022 · 11 citations
- On the efficiency of test suite based program repair: A Systematic Assessment of 16 Automated Repair Systems for Java ProgramsKui Liu, Shangwen Wang, Anil Koyuncu, Kisub Kim et al.ICSE 2020 · 116 citations
- A Dataset of Reproducible Flaky-Test FailuresSuzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan et al.ISSTA 2026
- A Large-Scale Empirical Review of Patch Correctness Checking ApproachesJun Yang, Yuehan Wang, Yiling Lou, Ming Wen et al.FSE 2023 · 11 citations
- SoK: Automated Vulnerability Repair: Methods, Tools, and AssessmentsYiwei Hu, Zhen Li, Kedie Shu, Shenghua Guan et al.USENIX Security 2025
