Aggregate Models, Not Explanations: Improving Feature Importance Estimation
Joseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion Bertrand
Abstract
Feature-importance methods show promise for transforming machine learning (ML) models from predictive engines into tools for scientific discovery. However, expressive models can be unstable due to data sampling and algorithmic stochasticity, leading to inaccurate variable importance estimates, undermining their utility in critical biomedical applications. While ensembling offers a remedy, the choice between explaining a single ensemble model or aggregating individual model explanations is non-trivial due to the non-linearity of importance measures, and remains largely understudied. Our theoretical analysis, developed under assumptions accommodating complex state-of-the-art ML models, reveals that this choice is governed by a trade-off involving the model's excess risk. In contrast to prior literature, we show that ensembling at the model level provides more accurate variable-importance estimates, particularly for expressive models, by reducing this leading error term. We validate these findings on classical benchmarks and a large-scale proteomic study from the UK Biobank.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79c6a050-f101-481e-871a-c2cdd3dc9f03Builds on7
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable ImportanceJon Donnelly, Srikar Katta, Cynthia Rudin, Edward P. BrowneNeurIPS 2023 · 41 citations
- Marginal Contribution Feature Importance - an Axiomatic Approach for Explaining DataAmnon Catav, Boyang Fu, Yazeed Zoabi, Ahuva Weiss-Meilik et al.ICML 2021 · 34 citations
- Statistically Valid Variable Importance Assessment through Conditional PermutationsAhmad Chamma, Denis A. Engemann, Bertrand ThirionNeurIPS 2023 · 23 citations
Related papers
- Towards a Unified Framework for Uncertainty-aware Nonlinear Variable Selection with Theoretical GuaranteesWenying Deng, Beau Coker, Rajarshi Mukherjee, Jeremiah Z. Liu et al.NeurIPS 2022 · 5 citations
- Variable Importance in High-Dimensional Settings Requires GroupingAhmad Chamma, Bertrand Thirion, Denis A. EngemannAAAI 2024 · 13 citations
- Implications of Model Indeterminacy for Explanations of Automated DecisionsMarc-Etienne Brunet, Ashton Anderson, Richard S. ZemelNeurIPS 2022 · 22 citations
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 115 citations
- Synthetic Model Combination: An Instance-wise Approach to Unsupervised Ensemble LearningAlex J. Chan, Mihaela van der SchaarNeurIPS 2022 · 6 citations
