Feature Importance Metrics in the Presence of Missing Data
Henrik von Kleist, Joshua Wendland, Ilya Shpitser, Carsten Marr
Abstract
Feature importance metrics are critical for interpreting machine learning models and understanding the relevance of individual features. However, real-world data often exhibit missingness, thereby complicating how feature importance should be evaluated. We introduce the distinction between two evaluation frameworks under missing data: (1) feature importance under the full data, as if every feature had been fully measured, and (2) feature importance under the observed data, where missingness is governed by the current measurement policy. While the full data perspective offers insights into the data generating process, it often relies on unrealistic assumptions and cannot guide decisions when missingness persists at model deployment. Since neither framework directly informs improvements in data collection, we additionally introduce the feature measurement importance gradient (FMIG), a novel, model-agnostic metric that identifies features that should be measured more frequently to enhance predictive performance. Using synthetic data, we illustrate key differences between these metrics and the risks of conflating them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- What's a good imputation to predict with missing values?Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël VaroquauxNeurIPS 2021 · 95 citations
- Full Law Identification in Graphical Models of Missing Data: Completeness ResultsRazieh Nabi, Rohit Bhattacharya, Ilya ShpitserICML 2020 · 60 citations
- Explaining Reinforcement Learning with Shapley ValuesDaniel Beechey, Thomas M. S. Smith, Özgür SimsekICML 2023 · 41 citations
Related papers
- Marginal Contribution Feature Importance - an Axiomatic Approach for Explaining DataAmnon Catav, Boyang Fu, Yazeed Zoabi, Ahuva Weiss-Meilik et al.ICML 2021 · 34 citations
- Prediction models that learn to avoid missing valuesLena Stempfle, Anton Matsson, Newton Mwai Kinyanjui, Fredrik D. JohanssonICML 2025
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Assessing Fairness in the Presence of Missing DataYiliang Zhang, Qi LongNeurIPS 2021 · 51 citations
- Interpretable Generalized Additive Models for Datasets with Missing ValuesHayden McTavish, Jon Donnelly, Margo I. Seltzer, Cynthia RudinNeurIPS 2024 · 9 citations
