Boosting for Predictive Sufficiency
Abbavaram Gowtham Reddy, Rajeev Verma, Celia Rubio-Madrigal, Krikamol Muandet, Rebekka Burkholz
摘要
Out-of-distribution (OOD) generalization is a defining hallmark of truly robust and reliable machine learning systems. Recently, it has been empirically observed that existing OOD generalization methods often underperform on real-world tabular data, where hidden confounding shifts drive distribution shifts that boosting models handle more effectively. Earlier work attributes a part of boosting's success to variance reduction, handling missing covariates, feature selection, and connections to multicalibration. Complementary to these explanations, we uncover a crucial reason behind boosting's success in OOD generalization: its ability to identify environments created by hidden confounding shifts and maximize predictive performance within those environments. To this end, this paper introduces an information-theoretic notion called α-predictive sufficiency and formalizes its connection to OOD generalization under hidden confounding shift. We show that boosting implicitly identifies suitable environments and produces an α-predictive sufficient predictor. We validate our theoretical results through synthetic and realworld experiments and show that boosting achieves robust performance by identifying these environments and maximizing the mutual information between predictions and true outcomes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann 等NeurIPS 2020 · 被引用 688 次
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 被引用 454 次
- Domain Adaptation with Conditional Distribution Matching and Generalized Label ShiftRemi Tachet des Combes, Han Zhao, Yu-Xiang Wang, Geoffrey J. GordonNeurIPS 2020 · 被引用 231 次
相关 Paper
- When Shift Happens - Confounding Is to BlameAbbavaram Gowtham Reddy, Celia Rubio-Madrigal, Rebekka Burkholz, Krikamol MuandetICLR 2026 · 被引用 5 次
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationKartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet 等NeurIPS 2021 · 被引用 372 次
- Bridging Multicalibration and Out-of-distribution Generalization Beyond Covariate ShiftJiayun Wu, Jiashuo Liu, Peng Cui, Steven WuNeurIPS 2024 · 被引用 14 次
- Explaining Concept Shift with Interpretable Feature AttributionRuiqi Lyu, Alistair Turcan, Bryan WilderICML 2026
- Overparameterization Improves Robustness to Covariate Shift in High DimensionsNilesh Tripuraneni, Ben Adlam, Jeffrey PenningtonNeurIPS 2021 · 被引用 50 次
