Efficient nonparametric statistical inference on population feature importance using Shapley values
Brian D. Williamson, Jean Feng
Abstract
The true population-level importance of a variable in a prediction task provides useful knowledge about the underlying data-generating mechanism and can help in deciding which measurements to collect in subsequent experiments. Valid statistical inference on this importance is a key component in understanding the population of interest. We present a computationally efficient procedure for estimating and obtaining valid statistical inference on the Shapley Population Variable Importance Measure (SPVIM). Although the computational complexity of the true SPVIM scales exponentially with the number of variables, we propose an estimator based on randomly sampling only Θ(n) feature subsets given n observations. We prove that our estimator converges at an asymptotically optimal rate. Moreover, by deriving the asymptotic distribution of our estimator, we construct valid confidence intervals and hypothesis tests. Our procedure has good finite-sample performance in simulations, and for an in-hospital mortality prediction task produces similar variable importance estimates when different machine learning algorithms are applied.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2031937-eba5-45bd-a005-47520becaa57Cited by top-tier papers19
- SHAP-IQ: Unified Approximation of any-order Shapley InteractionsFabian Fumagalli, Maximilian Muschalik, Patrick Kolpaczki, Eyke Hüllermeier et al.NeurIPS 2023 · 80 citations
- Measuring the Effect of Training Data on Deep Learning Predictions via Randomized ExperimentsJinkun Lin, Anqi Zhang, Mathias Lécuyer, Jinyang Li et al.ICML 2022 · 70 citations
- Explaining Predictive Uncertainty with Information Theoretic Shapley ValuesDavid S. Watson, Joshua O'Hara, Niek Tax, Richard Mudd et al.NeurIPS 2023 · 56 citations
- The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable ImportanceJon Donnelly, Srikar Katta, Cynthia Rudin, Edward P. BrowneNeurIPS 2023 · 41 citations
- Marginal Contribution Feature Importance - an Axiomatic Approach for Explaining DataAmnon Catav, Boyang Fu, Yazeed Zoabi, Ahuva Weiss-Meilik et al.ICML 2021 · 34 citations
Related papers
- Testing Conditional Mean Independence Using Generative Neural NetworksYi Zhang, Linjun Huang, Yun Yang, Xiaofeng ShaoICML 2025
- Measuring Variable Importance in Heterogeneous Treatment Effects with ConfidenceJoseph Paillard, Angel David Reyero Lobo, Vitaliy Kolodyazhniy, Bertrand Thirion et al.ICML 2025
- WeightedSHAP: analyzing and improving Shapley based feature attributionsYongchan Kwon, James Y. ZouNeurIPS 2022 · 60 citations
- Aggregate Models, Not Explanations: Improving Feature Importance EstimationJoseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion BertrandICML 2026 · 1 citation
- Lazy Estimation of Variable Importance for Large Neural NetworksYue Gao, Abby Stevens, Garvesh Raskutti, Rebecca WillettICML 2022 · 7 citations
