Out-Of-Domain Unlabeled Data Improves Generalization
Seyed Amir Hossein Saberi, Amir Najafi, Alireza Heidari, Mohammad Hosein Movasaghinia, Abolfazl S. Motahari, Babak H. Khalaj
Abstract
We propose a novel framework for incorporating unlabeled data into semi-supervised classification problems, where scenarios involving the minimization of either i) adversarially robust or ii) non-robust loss functions have been considered. Notably, we allow the unlabeled samples to deviate slightly (in total variation sense) from the in-domain distribution. The core idea behind our framework is to combine Distributionally Robust Optimization (DRO) with self-supervised training. As a result, we also leverage efficient polynomial-time algorithms for the training stage. From a theoretical standpoint, we apply our framework on the classification problem of a mixture of two Gaussians in , where in addition to the independent and labeled samples from the true distribution, a set of (usually with ) out of domain and unlabeled samples are given as well. Using only the labeled data, it is known that the generalization error can be bounded by . However, using our method on both isotropic and non-isotropic Gaussian mixture models, one can derive a new set of analytically explicit and non-asymptotic bounds which show substantial improvement on the generalization error compared to ERM. Our results underscore two significant insights: 1) out-of-domain samples, even when unlabeled, can be harnessed to narrow the generalization gap, provided that the true data distribution adheres to a form of the ``cluster assumption", and 2) the semi-supervised learning paradigm can be regarded as a special case of our framework when there are no distributional shifts. We validate our claims through experiments conducted on a variety of synthetic and real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14ca4a0b-ed30-40e8-ad5f-2177d88cd538Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Gradual Domain Adaptation via Manifold-Constrained Distributionally Robust OptimizationSeyed Amir Saberi, Amir Najafi, Amin Behjati, Ala Emrani et al.NeurIPS 2024 · 3 citations
- Complementary Benefits of Contrastive Learning and Self-Training Under Distribution ShiftSaurabh Garg, Amrith Setlur, Zachary C. Lipton, Sivaraman Balakrishnan et al.NeurIPS 2023 · 13 citations
- A Characterization of Semi-Supervised Adversarially Robust PAC LearnabilityIdan Attias, Steve Hanneke, Yishay MansourNeurIPS 2022 · 19 citations
- Generalized Semi-Supervised Learning via Self-Supervised Feature AdaptationJiachen Liang, Ruibing Hou, Hong Chang, Bingpeng Ma et al.NeurIPS 2023 · 7 citations
- Distributionally Robust Classification for Multi-source Unsupervised Domain AdaptationSeonghwi Kim, Sungho Jo, Wooseok Ha, Minwoo ChaeICLR 2026 · 4 citations
