Uncertainty in Gradient Boosting via Ensembles
Andrey Malinin, Liudmila Prokhorenkova, Aleksei Ustimenko
Abstract
Gradient boosting is a powerful machine learning technique that is particularly successful for tasks containing heterogeneous features and noisy data. While gradient boosting classification models return a distribution over class labels, regressions models typically yield only point predictions. However, for many practical, high-risk applications, it is also important to be able to quantify uncertainty in the predictions to avoid costly mistakes. In this work, we examine a probabilistic ensemble-based framework for deriving uncertainty estimates in the predictions of gradient boosting classification and regression models. Crucially, the proposed approach allows the total uncertainty to be decomposed into data uncertainty, which comes from the complexity and noise in data distribution, and knowledge uncertainty, coming from the lack of information about a given region of the feature space. Two approaches for generating ensembles are considered: Stochastic Gradient Boosting (SGB) and Stochastic Gradient Langevin Boosting (SGLB). Notably, SGLB also enables the generation of a virtual ensemble via only one gradient boosting model, which significantly reduces complexity. Experiments on a range of regression and classification datasets show that ensembles of gradient boosting models yield improved predictive performance, and measures of uncertainty successfully enable detection of out-of-domain inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 021a917f-0790-4da1-8823-75f52dde2afcCited by top-tier papers12
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas et al.ICML 2024 · 113 citations
- SGLB: Stochastic Gradient Langevin BoostingAleksei Ustimenko, Liudmila ProkhorenkovaICML 2021 · 20 citations
- Instance-Based Uncertainty Estimation for Gradient-Boosted Regression TreesJonathan Brophy, Daniel LowdNeurIPS 2022 · 17 citations
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- How does Bayesian Sampling help Membership Inference Attacks?Zhenlong Liu, Wenyu Jiang, Feng Zhou, Hongxin WeiICML 2026 · 3 citations
Builds on4
- NGBoost: Natural Gradient Boosting for Probabilistic PredictionTony Duan, Anand Avati, Daisy Yi Ding, Khanh K. Thai et al.ICML 2020 · 433 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- SGLB: Stochastic Gradient Langevin BoostingAleksei Ustimenko, Liudmila ProkhorenkovaICML 2021 · 20 citations
Related papers
- Wasserstein Gradient Boosting: A Framework for Distribution-Valued Supervised LearningTakuo MatsubaraNeurIPS 2024 · 7 citations
- Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic RegressionOlivier Sprangers, Sebastian Schelter, Maarten de RijkeKDD 2021 · 38 citations
- Smooth And Consistent Probabilistic Regression TreesSami Alkhoury, Emilie Devijver, Marianne Clausel, Myriam Tami et al.NeurIPS 2020 · 13 citations
- Learn Together Stop Apart: An Inclusive Approach to Ensemble PruningBulat Ibragimov, Gleb GusevKDD 2024 · 1 citation
- Learning Probabilistic Ordinal Embeddings for Uncertainty-Aware RegressionWanhua Li, Xiaoke Huang, Jiwen Lu, Jianjiang Feng et al.CVPR 2021
