RavenBuild: Context, Relevance, and Dependency Aware Build Outcome Prediction
Gengyi Sun, Sarra Habchi, Shane McIntosh
Abstract
Continuous Integration (CI) is a common practice adopted by modern software organizations. It plays an especially important role for large corporations like Ubisoft, where thousands of build jobs are submitted daily. Indeed, the cadence of development progress is constrained by the pace at which CI services process build jobs. To provide faster CI feedback, recent work explores how build outcomes can be anticipated. Although early results show plenty of promise, the distinct characteristics of Project X-a AAA video game project at Ubisoft-present new challenges for build outcome prediction. In the Project X setting, changes that do not modify source code also incur build failures. We also observe that the code changes that have an impact that crosses the source-data boundary are more prone to build failures than code changes that do not impact data files. Since such changes are not fully characterized by the existing set of features for build outcome prediction, state-of-the-art models tend to underperform.
To incorporate the data context, we propose RavenBuild-a novel approach to build outcome prediction that leverages context-, relevance-, and dependency-aware features. In the Project X context, we observe that RavenBuild improves the F1-score of the failing class by 50%, the recall of the failing class by 105%, and the AUC by 11% with respect to the state-of-the-art BuildFast approach. To ease adoption in settings with heterogeneous project sets, we also provide a simplified alternative RavenBuild-CR, which excludes dependency-aware features. We observe across-the-board improvements when RavenBuild-CR is applied to 22 open-source projects and Project X. On the other hand, we find that a naïve Parrot approach, which simply echoes the previous build outcome as its prediction, is surprisingly competitive with BuildFast and RavenBuild. Though Parrot fails to predict when the build outcome differs from their immediate predecessor, Parrot serves well as a tendency indicator of the sequences in build outcome datasets. Thus, we recommend that future studies also compare to the Parrot approach as a baseline when evaluating build outcome prediction models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7133c155-bb50-461f-bea7-08fb8f9425b9Cited by top-tier papers2
- Developer-Applied Accelerations in Continuous Integration: A Detection Approach and Catalog of PatternsMingyang Yin, Yutaro Kashiwa, Keheliya Gallaba, Mahmoud Alfadel et al.ASE 2024 · 5 citations
- Rechecking Recheck Requests in Continuous Integration: An Empirical Study of OpenStackYelizaveta Brus, Rungroj Maipradit, Earl T. Barr, Shane McIntoshASE 2025 · 1 citation
Builds on5
- A cost-efficient approach to building in continuous integrationXianhao Jin, Francisco ServantICSE 2020 · 34 citations
- BUILDFAST: History-Aware Build Outcome Prediction for Fast Feedback and Reduced Cost in Continuous IntegrationBihuan Chen, Linlin Chen, Chen Zhang, Xin PengASE 2020 · 32 citations
- Lessons from Eight Years of Operational Data from a Continuous Integration Service: An Exploratory Case Study of CircleCIKeheliya Gallaba, Maxime Lamothe, Shane McIntoshICSE 2022 · 28 citations
- Continuous test suite failure predictionCong Pan, Michael PradelISSTA 2021 · 22 citations
- Towards language-independent Brown Build DetectionDoriane Olewicki, Mathieu Nayrolles, Bram AdamsICSE 2022 · 17 citations
Related papers
- Commit Artifact Preserving Build PredictionGuoqing Wang, Zeyu Sun, Yizhou Chen, Yifan Zhao et al.ISSTA 2024 · 3 citations
- Understanding and Predicting Docker Build Duration: An Empirical Study of Containerized Workflow of OSS ProjectsYiwen Wu, Yang Zhang, Kele Xu, Tao Wang et al.ASE 2022 · 13 citations
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen et al.SIGMOD 2022 · 50 citations
- Escaping dependency hell: finding build dependency errors with the unified dependency graphGang Fan, Chengpeng Wang, Rongxin Wu, Xiao Xiao et al.ISSTA 2020 · 37 citations
- On the Benefits and Limits of Incremental Build of Software Configurations: An Exploratory StudyGeorges Aaron Randrianaina, Xhevahire Tërnava, Djamel Eddine Khelladi, Mathieu AcherICSE 2022 · 6 citations
