Debiasing Averaged Stochastic Gradient Descent to handle missing values
Aude Sportisse, Claire Boyer, Aymeric Dieuleveut, Julie Josse
Abstract
Stochastic gradient algorithm is a key ingredient of many machine learning methods, particularly appropriate for large-scale learning. However, a major caveat of large data is their incompleteness. We propose an averaged stochastic gradient algorithm handling missing values in linear models. This approach has the merit to be free from the need of any data distribution modeling and to account for heterogeneous missing proportion. In both streaming and finite-sample settings, we prove that this algorithm achieves convergence rate of O( 1 n ) at the iteration n, the same as without missing values. We show the convergence behavior and the relevance of the algorithm not only on synthetic data but also on real data sets, including those collected from medical register.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdd33e5e-3efd-4bb0-9ea6-f9cea8158a4dCited by top-tier papers3
- Conformal Prediction with Missing ValuesMargaux Zaffran, Aymeric Dieuleveut, Julie Josse, Yaniv RomanoICML 2023 · 31 citations
- Gradient Importance Learning for Incomplete ObservationsQitong Gao, Dong Wang, Joshua David Amason, Siyang Yuan et al.ICLR 2022 · 10 citations
- Near-optimal rate of consistency for linear models with missing valuesAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2022 · 10 citations
Builds on1
Related papers
- Accelerating SGD for Highly Ill-Conditioned Huge-Scale Online Matrix CompletionJialun Zhang, Hong-Ming Chiu, Richard Y. ZhangNeurIPS 2022 · 12 citations
- Random features models: a way to study the success of naive imputationAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2024 · 7 citations
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 10 citations
- Online Robust Regression via SGD on the l1 lossScott Pesme, Nicolas FlammarionNeurIPS 2020 · 41 citations
- Statistical Learning and Inverse Problems: A Stochastic Gradient ApproachYuri R. Fonseca, Yuri F. SaporitoNeurIPS 2022 · 7 citations
