Naive imputation implicitly regularizes high-dimensional linear models
Alexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan Scornet
摘要
Two different approaches exist to handle missing values for prediction: either imputation, prior to fitting any predictive algorithms, or dedicated methods able to natively incorporate missing values. While imputation is widely (and easily) use, it is unfortunately biased when low-capacity predictors (such as linear models) are applied afterward. However, in practice, naive imputation exhibits good predictive performance. In this paper, we study the impact of imputation in a high-dimensional linear model with MCAR missing data. We prove that zero imputation performs an implicit regularization closely related to the ridge method, often used in high-dimensional problems. Leveraging on this connection, we establish that the imputation bias is controlled by a ridge bias, which vanishes in high dimension. As a predictor, we argue in favor of the averaged SGD strategy, applied to zero-imputed data. We establish an upper bound on its generalization error, highlighting that imputation is benign in the d \sqrt n regime. Experiments illustrate our findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 被引用 179 次
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 135 次
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet 等NeurIPS 2020 · 被引用 50 次
- How to deal with missing data in supervised deep learning?Niels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2022 · 被引用 39 次
相关 Paper
- Random features models: a way to study the success of naive imputationAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2024 · 被引用 7 次
- Prediction models that learn to avoid missing valuesLena Stempfle, Anton Matsson, Newton Mwai Kinyanjui, Fredrik D. JohanssonICML 2025
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
- Why Not to Use Zero Imputation? Correcting Sparsity Bias in Training Neural NetworksJoonyoung Yi, Juhyuk Lee, Kwang Joon Kim, Sung Ju Hwang 等ICLR 2020 · 被引用 28 次
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 等NeurIPS 2021 · 被引用 41 次
