What's a good imputation to predict with missing values?
Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël Varoquaux
Abstract
How to learn a good predictor on data with missing values? Most efforts focus on first imputing as well as possible and second learning on the completed data to predict the outcome. Yet, this widespread practice has no theoretical grounding. Here we show that for almost all imputation functions, an impute-then-regress procedure with a powerful learner is Bayes optimal. This result holds for all missing-values mechanisms, in contrast with the classic statistical results that require missing-at-random settings to use imputation in probabilistic modeling. Moreover, it implies that perfect conditional imputation is not needed for good prediction asymptotically. In fact, we show that on perfectly imputed data the best regression function will generally be discontinuous, which makes it hard to learn. Crafting instead the imputation so as to leave the regression function unchanged simply shifts the problem to learning discontinuous imputations. Rather, we suggest that it is easier to learn imputation and regression jointly. We propose such a procedure, adapting NeuMiss, a neural network capturing the conditional links across observed and unobserved variables whatever the missing-value pattern. Experiments confirm that joint imputation and regression through NeuMiss is better than various two step procedures in our experiments with finite number of samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 168fd16f-0e1e-4057-b9d4-ab3efddbc46cCited by top-tier papers16
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
- Conformal Prediction with Missing ValuesMargaux Zaffran, Aymeric Dieuleveut, Julie Josse, Yaniv RomanoICML 2023 · 31 citations
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 20 citations
- TimeCHEAT: A Channel Harmony Strategy for Irregularly Sampled Multivariate Time Series AnalysisJiexi Liu, Meng Cao, Songcan ChenAAAI 2025 · 16 citations
- Sharing Pattern Submodels for Prediction with Missing ValuesLena Stempfle, Ashkan Panahi, Fredrik D. JohanssonAAAI 2023 · 9 citations
Builds on3
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet et al.NeurIPS 2020 · 50 citations
- Clairvoyance: A Pipeline Toolkit for Medical Time SeriesDaniel Jarrett, Jinsung Yoon, Ioana Bica, Zhaozhi Qian et al.ICLR 2021 · 43 citations
- How to deal with missing data in supervised deep learning?Niels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2022 · 39 citations
Related papers
- Inferring the Invisible: Neuro-Symbolic Rule Discovery for Missing Value ImputationWendi Ren, Ke Wan, Junyu Leng, Shuang LiICLR 2026
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 105 citations
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer et al.NeurIPS 2020 · 274 citations
- Probabilistic Imputation for Time-series Classification with Missing DataSeunghyun Kim, Hyunsu Kim, Eunggu Yun, Hwangrae Lee et al.ICML 2023 · 37 citations
- Learnable Prompt as Pseudo-Imputation: Rethinking the Necessity of Traditional EHR Data Imputation in Downstream Clinical PredictionWeibin Liao, Yinghao Zhu, Zhongji Zhang, Yuhang Wang et al.KDD 2025 · 2 citations
