Random features models: a way to study the success of naive imputation
Alexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan Scornet
摘要
Constant (naive) imputation is still widely used in practice as this is a first easy-to-use technique to deal with missing data. Yet, this simple method could be expected to induce a large bias for prediction purposes, as the imputed input may strongly differ from the true underlying data. However, recent works suggest that this bias is low in the context of high-dimensional linear predictors when data is supposed to be missing completely at random (MCAR). This paper completes the picture for linear predictors by confirming the intuition that the bias is negligible and that surprisingly naive imputation also remains relevant in very low dimension. To this aim, we consider a unique underlying random features model, which offers a rigorous framework for studying predictive performances, whilst the dimension of the observed features varies. Building on these theoretical results, we establish finite-sample bounds on stochastic gradient (SGD) predictors applied to zeroimputed data, a strategy particularly well suited for large-scale learning. If the MCAR assumption appears to be strong, we show that similar favorable behaviors occur for more complex missing data scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Random Feature Representation BoostingNikita Zozoulenko, Thomas Cass, Lukas GononICML 2025
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
相关 Paper
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 被引用 10 次
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 被引用 56 次
- On the Double Descent of Random Features Models Trained with SGDFanghui Liu, Johan A. K. Suykens, Volkan CevherNeurIPS 2022 · 被引用 11 次
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet 等NeurIPS 2020 · 被引用 50 次
- Debiasing Averaged Stochastic Gradient Descent to handle missing valuesAude Sportisse, Claire Boyer, Aymeric Dieuleveut, Julie JosseNeurIPS 2020 · 被引用 12 次
