Naive imputation implicitly regularizes high-dimensional linear models
Alexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan Scornet
Abstract
Two different approaches exist to handle missing values for prediction: either imputation, prior to fitting any predictive algorithms, or dedicated methods able to natively incorporate missing values. While imputation is widely (and easily) use, it is unfortunately biased when low-capacity predictors (such as linear models) are applied afterward. However, in practice, naive imputation exhibits good predictive performance. In this paper, we study the impact of imputation in a high-dimensional linear model with MCAR missing data. We prove that zero imputation performs an implicit regularization closely related to the ridge method, often used in high-dimensional problems. Leveraging on this connection, we establish that the imputation bias is controlled by a ridge bias, which vanishes in high dimension. As a predictor, we argue in favor of the averaged SGD strategy, applied to zero-imputed data. We establish an upper bound on its generalization error, highlighting that imputation is benign in the d \sqrt n regime. Experiments illustrate our findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c3c0eeb-fb4e-4324-baf2-de30e540de32Builds on4
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 179 citations
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 135 citations
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet et al.NeurIPS 2020 · 50 citations
- How to deal with missing data in supervised deep learning?Niels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2022 · 39 citations
Related papers
- Random features models: a way to study the success of naive imputationAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2024 · 7 citations
- Prediction models that learn to avoid missing valuesLena Stempfle, Anton Matsson, Newton Mwai Kinyanjui, Fredrik D. JohanssonICML 2025
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
- Why Not to Use Zero Imputation? Correcting Sparsity Bias in Training Neural NetworksJoonyoung Yi, Juhyuk Lee, Kwang Joon Kim, Sung Ju Hwang et al.ICLR 2020 · 28 citations
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu et al.NeurIPS 2021 · 41 citations
