Random features models: a way to study the success of naive imputation
Alexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan Scornet
Abstract
Constant (naive) imputation is still widely used in practice as this is a first easy-to-use technique to deal with missing data. Yet, this simple method could be expected to induce a large bias for prediction purposes, as the imputed input may strongly differ from the true underlying data. However, recent works suggest that this bias is low in the context of high-dimensional linear predictors when data is supposed to be missing completely at random (MCAR). This paper completes the picture for linear predictors by confirming the intuition that the bias is negligible and that surprisingly naive imputation also remains relevant in very low dimension. To this aim, we consider a unique underlying random features model, which offers a rigorous framework for studying predictive performances, whilst the dimension of the observed features varies. Building on these theoretical results, we establish finite-sample bounds on stochastic gradient (SGD) predictors applied to zeroimputed data, a strategy particularly well suited for large-scale learning. If the MCAR assumption appears to be strong, we show that similar favorable behaviors occur for more complex missing data scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Random Feature Representation BoostingNikita Zozoulenko, Thomas Cass, Lukas GononICML 2025
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
Related papers
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 10 citations
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 56 citations
- On the Double Descent of Random Features Models Trained with SGDFanghui Liu, Johan A. K. Suykens, Volkan CevherNeurIPS 2022 · 11 citations
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet et al.NeurIPS 2020 · 50 citations
- Debiasing Averaged Stochastic Gradient Descent to handle missing valuesAude Sportisse, Claire Boyer, Aymeric Dieuleveut, Julie JosseNeurIPS 2020 · 12 citations
