How to deal with missing data in supervised deep learning?
Niels Bruun Ipsen, Pierre-Alexandre Mattei, Jes Frellsen
Abstract
The issue of missing data in supervised learning has been largely overlooked, especially in the deep learning community. We investigate strategies to adapt neural architectures to handle missing values. Here, we focus on regression and classification problems where the features are assumed to be missing at random. Of particular interest are schemes that allow to reuse as-is a neural discriminative architecture. One scheme involves imputing the missing values with learnable constants. We propose a second novel approach that leverages recent advances in deep generative modelling. More precisely, a deep latent variable model can be learned jointly with the discriminative model, using importance-weighted variational inference in an end-to-end way. This hybrid approach, which mimics multiple imputation, also allows to impute the data, by relying on both the discriminative and the generative model. We also discuss ways of using a pre-trained generative model to train the discriminative one. In domains where powerful deep generative models are available, the hybrid approach leads to large performance gains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- What's a good imputation to predict with missing values?Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël VaroquauxNeurIPS 2021 · 95 citations
- not-MIWAE: Deep Generative Modelling with Missing not at Random DataNiels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2021 · 81 citations
- Probabilistic Imputation for Time-series Classification with Missing DataSeunghyun Kim, Hyunsu Kim, Eunggu Yun, Hwangrae Lee et al.ICML 2023 · 37 citations
- Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte CarloIgnacio Peis, Chao Ma, José Miguel Hernández-LobatoNeurIPS 2022 · 25 citations
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 10 citations
Builds on1
Related papers
- Variational Inference for Discriminative Learning with Generative Modeling of Feature IncompletionKohei Miyaguchi, Takayuki Katsuki, Akira Koseki, Toshiya IwamoriICLR 2022 · 4 citations
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 56 citations
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
- Missing Value Imputation on Multidimensional Time SeriesParikshit Bansal, Prathamesh Deshpande, Sunita SarawagiVLDB 2021 · 90 citations
- Incomplete Multi-view Deep Clustering with Data Imputation and AlignmentJiyuan Liu, Xinwang Liu, Xinhang Wan, Ke Liang et al.NeurIPS 2025 · 1 citation
