Statistical Inference for Gradient Boosting Regression
Haimo Fang, Kevin Tan, Giles Hooker
摘要
Gradient boosting is widely popular due to its flexibility and predictive accuracy. However, statistical inference and uncertainty quantification for gradient boosting remain challenging and under-explored. We propose a unified framework for statistical inference in gradient boosting regression. Our framework integrates dropout or parallel training with a recently proposed regularization procedure called Boulevard that allows for a central limit theorem (CLT) for boosting. With these enhancements, we surprisingly find that increasing the dropout rate and the number of trees grown in parallel at each iteration substantially enhances signal recovery and overall performance. Our resulting algorithms enjoy similar CLTs, which we use to construct built-in confidence intervals, prediction intervals, and rigorous hypothesis tests for assessing variable importance in only O(nd 2 ) time with the Nyström method. Numerical experiments verify the asymptotic normality and demonstrate that our algorithms perform well, do not require early stopping, interpolate between regularized boosting and random forests, and confirm the validity of their built-in statistical inference procedures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Classification with Valid and Adaptive CoverageYaniv Romano, Matteo Sesia, Emmanuel J. CandèsNeurIPS 2020 · 被引用 586 次
- NGBoost: Natural Gradient Boosting for Probabilistic PredictionTony Duan, Anand Avati, Daisy Yi Ding, Khanh K. Thai 等ICML 2020 · 被引用 433 次
- Uncertainty in Gradient Boosting via EnsemblesAndrey Malinin, Liudmila Prokhorenkova, Aleksei UstimenkoICLR 2021 · 被引用 117 次
- SGLB: Stochastic Gradient Langevin BoostingAleksei Ustimenko, Liudmila ProkhorenkovaICML 2021 · 被引用 20 次
相关 Paper
- Efficient nonparametric statistical inference on population feature importance using Shapley valuesBrian D. Williamson, Jean FengICML 2020 · 被引用 86 次
- The Implicit Delta MethodNathan Kallus, James McInerneyNeurIPS 2022 · 被引用 3 次
- Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical GuaranteesWenying Deng, Beau Coker, Rajarshi Mukherjee, Jeremiah Z. Liu 等NeurIPS 2022 · 被引用 5 次
- Instance-Based Uncertainty Estimation for Gradient-Boosted Regression TreesJonathan Brophy, Daniel LowdNeurIPS 2022 · 被引用 17 次
- Smaller, more accurate regression forests using tree alternating optimizationArman Zharmagambetov, Miguel Á. Carreira-PerpiñánICML 2020 · 被引用 34 次
