Regression Learning with Limited Observations of Multivariate Outcomes and Features
Yifan Sun, Grace Yi
Abstract
Multivariate linear regression models are broadly used to facilitate relationships between outcomes and features. However, their effectiveness is compromised by the presence of missing observations, a ubiquitous challenge in real-world applications. Considering a scenario where learners access only limited components for both outcomes and features, we develop efficient algorithms tailored for the least squares (L 2 ) and least absolute (L 1 ) loss functions, each coupled with a ridge-like and Lasso-type penalty, respectively. Moreover, we establish rigorous error bounds for all proposed algorithms. Notably, our L 2 loss function algorithms are probably approximately correct (PAC), distinguishing them from their L 1 counterparts. Extensive numerical experiments show that our approach outperforms methods that apply existing algorithms for univariate outcome individually to each coordinate of multivariate outcomes in a naive manner. Further, utilizing the L 1 loss function or introducing a Lasso-type penalty can enhance predictions in the presence of outliers or high dimensional features. This research contributes valuable insights into addressing the challenges posed by incomplete data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Near-optimal rate of consistency for linear models with missing valuesAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2022 · 10 citations
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 10 citations
- Random features models: a way to study the success of naive imputationAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2024 · 7 citations
- Single Point Transductive PredictionNilesh Tripuraneni, Lester MackeyICML 2020 · 5 citations
- Can we globally optimize cross-validation loss? Quasiconvexity in ridge regressionWilliam T. Stephenson, Zachary Frangella, Madeleine Udell, Tamara BroderickNeurIPS 2021 · 15 citations
