Aggregate Models, Not Explanations: Improving Feature Importance Estimation
Joseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion Bertrand
摘要
Feature-importance methods show promise for transforming machine learning (ML) models from predictive engines into tools for scientific discovery. However, expressive models can be unstable due to data sampling and algorithmic stochasticity, leading to inaccurate variable importance estimates, undermining their utility in critical biomedical applications. While ensembling offers a remedy, the choice between explaining a single ensemble model or aggregating individual model explanations is non-trivial due to the non-linearity of importance measures, and remains largely understudied. Our theoretical analysis, developed under assumptions accommodating complex state-of-the-art ML models, reveals that this choice is governed by a trade-off involving the model's excess risk. In contrast to prior literature, we show that ensembling at the model level provides more accurate variable-importance estimates, particularly for expressive models, by reducing this leading error term. We validate these findings on classical benchmarks and a large-scale proteomic study from the UK Biobank.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable ImportanceJon Donnelly, Srikar Katta, Cynthia Rudin, Edward P. BrowneNeurIPS 2023 · 被引用 41 次
- Marginal Contribution Feature Importance - an Axiomatic Approach for Explaining DataAmnon Catav, Boyang Fu, Yazeed Zoabi, Ahuva Weiss-Meilik 等ICML 2021 · 被引用 34 次
- Statistically Valid Variable Importance Assessment through Conditional PermutationsAhmad Chamma, Denis A. Engemann, Bertrand ThirionNeurIPS 2023 · 被引用 23 次
相关 Paper
- Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical GuaranteesWenying Deng, Beau Coker, Rajarshi Mukherjee, Jeremiah Z. Liu 等NeurIPS 2022 · 被引用 5 次
- Variable Importance in High-Dimensional Settings Requires GroupingAhmad Chamma, Bertrand Thirion, Denis A. EngemannAAAI 2024 · 被引用 13 次
- Implications of Model Indeterminacy for Explanations of Automated DecisionsMarc-Etienne Brunet, Ashton Anderson, Richard S. ZemelNeurIPS 2022 · 被引用 22 次
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 被引用 115 次
- Synthetic Model Combination: An Instance-wise Approach to Unsupervised Ensemble LearningAlex J. Chan, Mihaela van der SchaarNeurIPS 2022 · 被引用 6 次
