What's a good imputation to predict with missing values?
Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël Varoquaux
摘要
How to learn a good predictor on data with missing values? Most efforts focus on first imputing as well as possible and second learning on the completed data to predict the outcome. Yet, this widespread practice has no theoretical grounding. Here we show that for almost all imputation functions, an impute-then-regress procedure with a powerful learner is Bayes optimal. This result holds for all missing-values mechanisms, in contrast with the classic statistical results that require missing-at-random settings to use imputation in probabilistic modeling. Moreover, it implies that perfect conditional imputation is not needed for good prediction asymptotically. In fact, we show that on perfectly imputed data the best regression function will generally be discontinuous, which makes it hard to learn. Crafting instead the imputation so as to leave the regression function unchanged simply shifts the problem to learning discontinuous imputations. Rather, we suggest that it is easier to learn imputation and regression jointly. We propose such a procedure, adapting NeuMiss, a neural network capturing the conditional links across observed and unobserved variables whatever the missing-value pattern. Experiments confirm that joint imputation and regression through NeuMiss is better than various two step procedures in our experiments with finite number of samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth 等ICML 2022 · 被引用 129 次
- Conformal Prediction with Missing ValuesMargaux Zaffran, Aymeric Dieuleveut, Julie Josse, Yaniv RomanoICML 2023 · 被引用 31 次
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 被引用 20 次
- TimeCHEAT: A Channel Harmony Strategy for Irregularly Sampled Multivariate Time Series AnalysisJiexi Liu, Meng Cao, Songcan ChenAAAI 2025 · 被引用 16 次
- Sharing Pattern Submodels for Prediction with Missing ValuesLena Stempfle, Ashkan Panahi, Fredrik D. JohanssonAAAI 2023 · 被引用 9 次
它引用的顶会 Paper3
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet 等NeurIPS 2020 · 被引用 50 次
- Clairvoyance: A Pipeline Toolkit for Medical Time SeriesDaniel Jarrett, Jinsung Yoon, Ioana Bica, Zhaozhi Qian 等ICLR 2021 · 被引用 43 次
- How to deal with missing data in supervised deep learning?Niels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2022 · 被引用 39 次
相关 Paper
- Inferring the Invisible: Neuro-Symbolic Rule Discovery for Missing Value ImputationWendi Ren, Ke Wan, Junyu Leng, Shuang LiICLR 2026
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 被引用 105 次
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer 等NeurIPS 2020 · 被引用 274 次
- Probabilistic Imputation for Time-series Classification with Missing DataSeunghyun Kim, Hyunsu Kim, Eunggu Yun, Hwangrae Lee 等ICML 2023 · 被引用 37 次
- Learnable Prompt as Pseudo-Imputation: Rethinking the Necessity of Traditional EHR Data Imputation in Downstream Clinical PredictionWeibin Liao, Yinghao Zhu, Zhongji Zhang, Yuhang Wang 等KDD 2025 · 被引用 2 次
