Near-optimal rate of consistency for linear models with missing values
Alexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan Scornet
Abstract
Missing values arise in most real-world data sets due to the aggregation of multiple sources and intrinsically missing information (sensor failure, unanswered questions in surveys...). In fact, the very nature of missing values usually prevents us from running standard learning algorithms. In this paper, we focus on the extensively-studied linear models, but in presence of missing values, which turns out to be quite a challenging task. Indeed, the Bayes predictor can be decomposed as a sum of predictors corresponding to each missing pattern. This eventually requires to solve a number of learning tasks, exponential in the number of input features, which makes predictions impossible for current real-world datasets. First, we propose a rigorous setting to analyze a leastsquare type estimator and establish a bound on the excess risk which increases exponentially in the dimension. Consequently, we leverage the missing data distribution to propose a new algorithm, and derive associated adaptive risk bounds that turn out to be minimax optimal. Numerical experiments highlight the benefits of our method compared to state-of-the-art algorithms used for predictions with missing values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on2
- NeuMiss networks: differentiable programming for supervised learning with missing valuesMarine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet et al.NeurIPS 2020 · 50 citations
- Debiasing Averaged Stochastic Gradient Descent to handle missing valuesAude Sportisse, Claire Boyer, Aymeric Dieuleveut, Julie JosseNeurIPS 2020 · 12 citations
Related papers
- Regression Learning with Limited Observations of Multivariate Outcomes and FeaturesYifan Sun, Grace YiICML 2024
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 45 citations
- On Learning Mixture of Linear Regressions in the Non-Realizable SettingSoumyabrata Pal, Arya Mazumdar, Rajat Sen, Avishek GhoshICML 2022 · 13 citations
- Random features models: a way to study the success of naive imputationAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2024 · 7 citations
- Variational Inference for Discriminative Learning with Generative Modeling of Feature IncompletionKohei Miyaguchi, Takayuki Katsuki, Akira Koseki, Toshiya IwamoriICLR 2022 · 4 citations
