Wasserstein Gradient Boosting: A Framework for Distribution-Valued Supervised Learning
Takuo Matsubara
Abstract
Gradient boosting is a sequential ensemble method that fits a new weaker learner to pseudo residuals at each iteration. We propose Wasserstein gradient boosting, a novel extension of gradient boosting that fits a new weak learner to alternative pseudo residuals that are Wasserstein gradients of loss functionals of probability distributions assigned at each input. It solves distribution-valued supervised learning, where the output values of the training dataset are probability distributions for each input. In classification and regression, a model typically returns, for each input, a point estimate of a parameter of a noise distribution specified for a response variable, such as the class probability parameter of a categorical distribution specified for a response label. A main application of Wasserstein gradient boosting in this paper is tree-based evidential learning, which returns a distributional estimate of the response parameter for each input. We empirically demonstrate the superior performance of the probabilistic prediction by Wasserstein gradient boosting in comparison with existing uncertainty quantification methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4cb3b18-7dfa-436b-b27f-37fb72753b6bCited by top-tier papers2
- Fréchet Geodesic BoostingYidong Zhou, Su I Iao, Hans-Georg MüllerNeurIPS 2025
- GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category DiscoveryXi Chen, Chuan Qin, Jinpeng Li, Shasha Hu et al.ACL 2026
Builds on6
- Deep Evidential RegressionAlexander Amini, Wilko Schwarting, Ava Soleimany, Daniela RusNeurIPS 2020 · 777 citations
- NGBoost: Natural Gradient Boosting for Probabilistic PredictionTony Duan, Anand Avati, Daisy Yi Ding, Khanh K. Thai et al.ICML 2020 · 433 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-CountsBertrand Charpentier, Daniel Zügner, Stephan GünnemannNeurIPS 2020 · 263 citations
- A Non-Asymptotic Analysis for Stein Variational Gradient DescentAnna Korba, Adil Salim, Michael Arbel, Giulia Luise et al.NeurIPS 2020 · 102 citations
Related papers
- Uncertainty in Gradient Boosting via EnsemblesAndrey Malinin, Liudmila Prokhorenkova, Aleksei UstimenkoICLR 2021 · 117 citations
- Treeffuser: probabilistic prediction via conditional diffusions with gradient-boosted treesNicolas Beltran-Velez, Alessandro Antonio Grande, Achille Nazaret, Alp Kucukelbir et al.NeurIPS 2024 · 8 citations
- Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic RegressionOlivier Sprangers, Sebastian Schelter, Maarten de RijkeKDD 2021 · 38 citations
- Instance-Based Uncertainty Estimation for Gradient-Boosted Regression TreesJonathan Brophy, Daniel LowdNeurIPS 2022 · 17 citations
- Smooth And Consistent Probabilistic Regression TreesSami Alkhoury, Emilie Devijver, Marianne Clausel, Myriam Tami et al.NeurIPS 2020 · 13 citations
