Semi-supervised Active Linear Regression
Nived Rajaraman, Devvrit, Pranjal Awasthi
Abstract
Labeled data often comes at a high cost as it may require recruiting human labelers or running costly' experiments. At the same time, in many practical scenarios, one already has access to a partially labeled, potentially biased dataset that can help with the learning task at hand. Motivated by such settings, we formally initiate a study of semi-supervised active learning through the frame of linear regression. Here, the learner has access to a dataset X ∈ R (nun+nlab)×d composed of n un unlabeled examples that a learner can actively query, and n lab examples labeled a priori. Denoting the true labels by Y ∈ R nun+nlab , the learner's objective is to find while querying the labels of as few unlabeled points as possible. In this paper, we introduce an instance dependent parameter called the reduced rank, denoted R X , and propose an efficient algorithm with query complexity O(R X / ). This result directly implies improved upper bounds for two important special cases: (i) active ridge regression, and (ii) active kernel ridge regression, where the reducedrank equates to the statistical dimension, sd λ and effective dimension, d λ of the problem respectively, where λ ≥ 0 denotes the regularization parameter. Finally, we introduce a distributional version of the problem as a special case of the agnostic formulation we consider earlier; here, for every X, we prove a matching instancewise lower bound of Ω(R X / ) on the query complexity of any algorithm. * equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Semi-Supervised Learning with Noisy Proxy Covariates: Generalization Bounds and Distribution RegressionKwangho Kim, Jisu KimICML 2026
- Distributionally Robust Active Learning for Gaussian Process RegressionShion Takeno, Yoshito Okura, Yu Inatsu, Tatsuya Aoyama et al.ICML 2025
- Active Labeling: Streaming Stochastic GradientsVivien Cabannes, Francis R. Bach, Vianney Perchet, Alessandro RudiNeurIPS 2022 · 2 citations
- Active Learning of General Halfspaces: Label Queries vs Membership QueriesIlias Diakonikolas, Daniel M. Kane, Mingchen MaNeurIPS 2024 · 7 citations
- A Neural Pre-Conditioning Active Learning Algorithm to Reduce Label ComplexitySeo Taek Kong, Soomin Jeon, Dongbin Na, Jaewon Lee et al.NeurIPS 2022 · 7 citations
