Towards language-independent Brown Build Detection
Doriane Olewicki, Mathieu Nayrolles, Bram Adams
Abstract
In principle, continuous integration (CI) practices allow modern software organizations to build and test their products after each code change to detect quality issues as soon as possible. In reality, issues with the build scripts (e.g., missing dependencies) and/or the presence of "flaky tests" lead to build failures that essentially are false positives, not indicative of actual quality problems of the source code. For our industrial partner, which is active in the video game industry, such "brown builds" not only require multidisciplinary teams to spend more effort interpreting or even re-running the build, leading to substantial redundant build activity, but also slows down the integration pipeline. Hence, this paper aims to prototype and evaluate approaches for early detection of brown build results based on textual similarity to build logs of prior brown builds. The approach is tested on 7 projects (6 closed-source from our industrial collaborators and 1 open-source, Graphviz). We find that our model manages to detect brown builds with a mean F1-score of 53% on the studied projects, which is three times more than the best baseline considered, and at least as good as human experts (but with less effort). Furthermore, we found that cross-project prediction can be used for a project's onboarding phase, that a training set of 30-weeks works best, and that our retraining heuristics keep the F1-score higher than the baseline, while retraining only every 4--5 weeks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- RavenBuild: Context, Relevance, and Dependency Aware Build Outcome PredictionGengyi Sun, Sarra Habchi, Shane McIntoshFSE 2024 · 8 citations
- Repeated Builds During Code Review: An Empirical Study of the OpenStack CommunityRungroj Maipradit, Dong Wang, Patanamon Thongtanunam, Raula Gaikovina Kula et al.ASE 2023 · 8 citations
- An Empirical Study on Code Review Activity Prediction and Its Impact in PracticeDoriane Olewicki, Sarra Habchi, Bram AdamsFSE 2024 · 2 citations
- Rechecking Recheck Requests in Continuous Integration: An Empirical Study of OpenStackYelizaveta Brus, Rungroj Maipradit, Earl T. Barr, Shane McIntoshASE 2025 · 1 citation
Related papers
- What Happened in This Pipeline? Diffing Build Logs with CiDiffNicolas Hubner, Jean-Rémy Falleri, Raluca Uricaru, Thomas Degueule et al.ISSTA 2025
- BUILDFAST: History-Aware Build Outcome Prediction for Fast Feedback and Reduced Cost in Continuous IntegrationBihuan Chen, Linlin Chen, Chen Zhang, Xin PengASE 2020 · 32 citations
- Commit Artifact Preserving Build PredictionGuoqing Wang, Zeyu Sun, Yizhou Chen, Yifan Zhao et al.ISSTA 2024 · 3 citations
- Buildsheriff: Change-Aware Test Failure Triage for Continuous Integration BuildsChen Zhang, Bihuan Chen, Xin Peng, Wenyun ZhaoICSE 2022 · 10 citations
- Escaping dependency hell: finding build dependency errors with the unified dependency graphGang Fan, Chengpeng Wang, Rongxin Wu, Xiao Xiao et al.ISSTA 2020 · 37 citations
