Uncertainty in Gradient Boosting via Ensembles
Andrey Malinin, Liudmila Prokhorenkova, Aleksei Ustimenko
摘要
Gradient boosting is a powerful machine learning technique that is particularly successful for tasks containing heterogeneous features and noisy data. While gradient boosting classification models return a distribution over class labels, regressions models typically yield only point predictions. However, for many practical, high-risk applications, it is also important to be able to quantify uncertainty in the predictions to avoid costly mistakes. In this work, we examine a probabilistic ensemble-based framework for deriving uncertainty estimates in the predictions of gradient boosting classification and regression models. Crucially, the proposed approach allows the total uncertainty to be decomposed into data uncertainty, which comes from the complexity and noise in data distribution, and knowledge uncertainty, coming from the lack of information about a given region of the feature space. Two approaches for generating ensembles are considered: Stochastic Gradient Boosting (SGB) and Stochastic Gradient Langevin Boosting (SGLB). Notably, SGLB also enables the generation of a virtual ensemble via only one gradient boosting model, which significantly reduces complexity. Experiments on a range of regression and classification datasets show that ensembles of gradient boosting models yield improved predictive performance, and measures of uncertainty successfully enable detection of out-of-domain inputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas 等ICML 2024 · 被引用 113 次
- SGLB: Stochastic Gradient Langevin BoostingAleksei Ustimenko, Liudmila ProkhorenkovaICML 2021 · 被引用 20 次
- Instance-Based Uncertainty Estimation for Gradient-Boosted Regression TreesJonathan Brophy, Daniel LowdNeurIPS 2022 · 被引用 17 次
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- How does Bayesian Sampling help Membership Inference Attacks?Zhenlong Liu, Wenyu Jiang, Feng Zhou, Hongxin WeiICML 2026 · 被引用 3 次
它引用的顶会 Paper4
- NGBoost: Natural Gradient Boosting for Probabilistic PredictionTony Duan, Anand Avati, Daisy Yi Ding, Khanh K. Thai 等ICML 2020 · 被引用 433 次
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 被引用 354 次
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 被引用 273 次
- SGLB: Stochastic Gradient Langevin BoostingAleksei Ustimenko, Liudmila ProkhorenkovaICML 2021 · 被引用 20 次
相关 Paper
- Wasserstein Gradient Boosting: A Framework for Distribution-Valued Supervised LearningTakuo MatsubaraNeurIPS 2024 · 被引用 7 次
- Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic RegressionOlivier Sprangers, Sebastian Schelter, Maarten de RijkeKDD 2021 · 被引用 38 次
- Smooth And Consistent Probabilistic Regression TreesSami Alkhoury, Emilie Devijver, Marianne Clausel, Myriam Tami 等NeurIPS 2020 · 被引用 13 次
- Learn Together Stop Apart: An Inclusive Approach to Ensemble PruningBulat Ibragimov, Gleb GusevKDD 2024 · 被引用 1 次
- Learning Probabilistic Ordinal Embeddings for Uncertainty-Aware RegressionWanhua Li, Xiaoke Huang, Jiwen Lu, Jianjiang Feng 等CVPR 2021
