CoffeeBoost: Gradient Boosting Native Conformal Inference for Bayesian Optimization
Yuanhao Lai, Pengfei Zheng, Chenpeng Ji, Cheng Qiu, Tingkai Wang, Songhan Zhang, Zhengang Wang, Yunfei Du
Abstract
Bayesian optimization (BO) is a key technique for solving black-box optimization problems. This study extends the scope of BO from conventional applications (e.g., AutoML and robotics learning) to automated tuning of software systems. Despite GP (Gaussian Process) implementing a foundation formalism for exploitation and exploration in BO, its limited predictive power and unrealistic assumptions (e.g., continuity and Gaussianity) can severely affect its effectiveness and efficiency in tuning complex software systems. To overcome these limitations, we propose a BO framework Coffee-Boost, which implements exploitation and exploration with a GBDT-native distribution-free probabilistic surrogate model. CoffeeBoost constructs surrogate models via stochastic gradient boosting ensembles (SGBE) and quantifies probabilistic distributions via distribution-free conformal predictive systems. Moreover, CoffeeBoost leverages the residual paths in SGBE to improve the local adaptiveness of the resulting predictive distributions in a GBDT-native manner. Across eight auto-tuning benchmarks for database management systems (DBMS), we evaluate CoffeeBoost and show its superior learnability and optimizability against existing GP-based and tree-ensemble-based BO schemes. Detailed analysis further shows CoffeeBoost's predictive distributions excel in both coverage and tightness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- NGBoost: Natural Gradient Boosting for Probabilistic PredictionTony Duan, Anand Avati, Daisy Yi Ding, Khanh K. Thai et al.ICML 2020 · 433 citations
- BANANAS: Bayesian Optimization with Neural Architectures for Neural Architecture SearchColin White, Willie Neiswanger, Yash SavaniAAAI 2021 · 401 citations
- Unexpected Improvements to Expected Improvement for Bayesian OptimizationSebastian Ament, Samuel Daulton, David Eriksson, Maximilian Balandat et al.NeurIPS 2023 · 280 citations
- Conformal prediction interval for dynamic time-seriesChen Xu, Yao XieICML 2021 · 174 citations
- Uncertainty in Gradient Boosting via EnsemblesAndrey Malinin, Liudmila Prokhorenkova, Aleksei UstimenkoICLR 2021 · 117 citations
Related papers
- Centrum: Model-based Database Auto-tuning with Minimal Distributional AssumptionsYuanhao Lai, Pengfei Zheng, Chenpeng Ji, Yan Li et al.SIGMOD 2025 · 1 citation
- Do the Best Cloud Configurations Grow on Trees? An Experimental Evaluation of Black Box Algorithms for Optimizing Cloud Workloads SubMuhammad Bilal, Marco Serafini, Marco Canini, Rodrigo RodriguesVLDB 2020
- Knowing The What But Not The Where in Bayesian OptimizationVu Nguyen, Michael A. OsborneICML 2020 · 42 citations
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 143 citations
- DivBO: Diversity-aware CASH for Ensemble LearningYu Shen, Yupeng Lu, Yang Li, Yaofeng Tu et al.NeurIPS 2022 · 15 citations
