Prediction models that learn to avoid missing values
Lena Stempfle, Anton Matsson, Newton Mwai Kinyanjui, Fredrik D. Johansson
摘要
Handling missing values at test time is challenging for machine learning models, especially when aiming for both high accuracy and interpretability. Established approaches often add bias through imputation or excessive model complexity via missingness indicators. Moreover, either method can obscure interpretability, making it harder to understand how the model utilizes the observed variables in predictions. We propose missingness-avoiding (MA) machine learning, a general framework for training models to rarely require the values of missing (or imputed) features at test time. We create tailored MA learning algorithms for decision trees, tree ensembles, and sparse linear models by incorporating classifierspecific regularization terms in their learning objectives. The tree-based models leverage contextual missingness by reducing reliance on missing values based on the observed context. Experiments on real-world datasets demonstrate that MA-DT, MA-LASSO, MA-RF, and MA-GBT effectively reduce the reliance on features with missing values while maintaining predictive performance competitive with their unregularized counterparts. This shows that our framework gives practitioners a powerful tool to maintain interpretability in predictions with test-time missing values.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 被引用 105 次
- What's a good imputation to predict with missing values?Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël VaroquauxNeurIPS 2021 · 被引用 95 次
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet 等NeurIPS 2020 · 被引用 50 次
- Sharing Pattern Submodels for Prediction with Missing ValuesLena Stempfle, Ashkan Panahi, Fredrik D. JohanssonAAAI 2023 · 被引用 9 次
- Interpretable Generalized Additive Models for Datasets with Missing ValuesHayden McTavish, Jon Donnelly, Margo I. Seltzer, Cynthia RudinNeurIPS 2024 · 被引用 9 次
相关 Paper
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 被引用 10 次
- Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing ValuesHaewon Jeong, Hao Wang, Flávio P. CalmonAAAI 2022 · 被引用 48 次
- Certain and Approximately Certain Models for Statistical LearningCheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh ChabadaSIGMOD 2024 · 被引用 4 次
- Feature Importance Metrics in the Presence of Missing DataHenrik von Kleist, Joshua Wendland, Ilya Shpitser, Carsten MarrICML 2025
- Leveraging Predictive Equivalence in Decision TreesHayden McTavish, Zachery Boner, Jon Donnelly, Margo I. Seltzer 等ICML 2025
