Interpretable Generalized Additive Models for Datasets with Missing Values
Hayden McTavish, Jon Donnelly, Margo I. Seltzer, Cynthia Rudin
Abstract
Many important datasets contain samples that are missing one or more feature values. Maintaining the interpretability of machine learning models in the presence of such missing data is challenging. Singly or multiply imputing missing values complicates the model's mapping from features to labels. On the other hand, reasoning on indicator variables that represent missingness introduces a potentially large number of additional terms, sacrificing sparsity. We solve these problems with M-GAM, a sparse, generalized, additive modeling approach that incorporates missingness indicators and their interaction terms while maintaining sparsity through l0 regularization. We show that M-GAM provides similar or superior accuracy to prior methods while significantly improving sparsity relative to either imputation or naive inclusion of indicator variables.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9964dad-b364-4556-bc77-bbebd87a3c30Cited by top-tier papers4
- Additive Models Explained: A Computational Complexity ApproachShahaf Bassan, Michal Moshkovitz, Guy KatzNeurIPS 2025 · 4 citations
- KANFIS: A Neuro-Symbolic Framework for Interpretable and Uncertainty-Aware LearningBinbin Yong, Haoran Pei, Jun Shen, Haoran Li et al.ICML 2026 · 2 citations
- Prediction models that learn to avoid missing valuesLena Stempfle, Anton Matsson, Newton Mwai Kinyanjui, Fredrik D. JohanssonICML 2025
- Leveraging Predictive Equivalence in Decision TreesHayden McTavish, Zachery Boner, Jon Donnelly, Margo I. Seltzer et al.ICML 2025
Builds on2
Related papers
- How Interpretable and Trustworthy are GAMs?Chun-Hao Chang, Sarah Tan, Benjamin J. Lengerich, Anna Goldenberg et al.KDD 2021 · 60 citations
- pureGAM: Learning an Inherently Pure Additive ModelXingzhi Sun, Ziyu Wang, Rui Ding, Shi Han et al.KDD 2022 · 4 citations
- GRAND-SLAMIN' Interpretable Additive Modeling with Structural ConstraintsShibal Ibrahim, Gabriel Afriat, Kayhan Behdin, Rahul MazumderNeurIPS 2023 · 15 citations
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 10 citations
