Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic Regression
Olivier Sprangers, Sebastian Schelter, Maarten de Rijke
Abstract
Gradient Boosting Machines (GBMs) are hugely popular for solving tabular data problems. However, practitioners are not only interested in point predictions, but also in probabilistic predictions in order to quantify the uncertainty of the predictions. Creating such probabilistic predictions is difficult with existing GBM-based solutions: they either require training multiple models or they become too computationally expensive to be useful for large-scale settings. We propose Probabilistic Gradient Boosting Machines (PGBMs), a method to create probabilistic predictions with a single ensemble of decision trees in a computationally efficient manner. PGBM approximates the leaf weights in a decision tree as a random variable, and approximates the mean and variance of each sample in a dataset via stochastic tree ensemble update equations. These learned moments allow us to subsequently sample from a specified distribution after training. We empirically demonstrate the advantages of PGBM compared to existing state-of-the-art methods: (i) PGBM enables probabilistic estimates without compromising on point performance in a single model, (ii) PGBM learns probabilistic estimates via a single model only (and without requiring multi-parameter boosting), and thereby offers a speedup of up to several orders of magnitude over existing state-of-the-art methods on large datasets, and (iii) PGBM achieves accurate probabilistic estimates in tasks with complex differentiable loss functions, such as hierarchical time series problems, where we observed up to 10% improvement in point forecasting performance and up to 300% improvement in probabilistic forecasting performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3028b867-2477-462c-932f-8d546e7b956dCited by top-tier papers6
- Instance-Based Uncertainty Estimation for Gradient-Boosted Regression TreesJonathan Brophy, Daniel LowdNeurIPS 2022 · 17 citations
- Probabilistic Forecasting: A Level-Set ApproachHilaf Hasson, Bernie Wang, Tim Januschowski, Jan GasthausNeurIPS 2021 · 15 citations
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Treeffuser: probabilistic prediction via conditional diffusions with gradient-boosted treesNicolas Beltran-Velez, Alessandro Antonio Grande, Achille Nazaret, Alp Kucukelbir et al.NeurIPS 2024 · 8 citations
- CoffeeBoost: Gradient Boosting Native Conformal Inference for Bayesian OptimizationYuanhao Lai, Pengfei Zheng, Chenpeng Ji, Cheng Qiu et al.AAAI 2025 · 1 citation
Builds on1
Related papers
- Uncertainty in Gradient Boosting via EnsemblesAndrey Malinin, Liudmila Prokhorenkova, Aleksei UstimenkoICLR 2021 · 117 citations
- Wasserstein Gradient Boosting: A Framework for Distribution-Valued Supervised LearningTakuo MatsubaraNeurIPS 2024 · 7 citations
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 407 citations
- SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput ProblemsLeonid Iosipoi, Anton VakhrushevNeurIPS 2022 · 20 citations
- Smooth And Consistent Probabilistic Regression TreesSami Alkhoury, Emilie Devijver, Marianne Clausel, Myriam Tami et al.NeurIPS 2020 · 13 citations
