Out-Of-Domain Unlabeled Data Improves Generalization
Seyed Amir Hossein Saberi, Amir Najafi, Alireza Heidari, Mohammad Hosein Movasaghinia, Abolfazl S. Motahari, Babak H. Khalaj
摘要
We propose a novel framework for incorporating unlabeled data into semi-supervised classification problems, where scenarios involving the minimization of either i) adversarially robust or ii) non-robust loss functions have been considered. Notably, we allow the unlabeled samples to deviate slightly (in total variation sense) from the in-domain distribution. The core idea behind our framework is to combine Distributionally Robust Optimization (DRO) with self-supervised training. As a result, we also leverage efficient polynomial-time algorithms for the training stage. From a theoretical standpoint, we apply our framework on the classification problem of a mixture of two Gaussians in , where in addition to the independent and labeled samples from the true distribution, a set of (usually with ) out of domain and unlabeled samples are given as well. Using only the labeled data, it is known that the generalization error can be bounded by . However, using our method on both isotropic and non-isotropic Gaussian mixture models, one can derive a new set of analytically explicit and non-asymptotic bounds which show substantial improvement on the generalization error compared to ERM. Our results underscore two significant insights: 1) out-of-domain samples, even when unlabeled, can be harnessed to narrow the generalization gap, provided that the true data distribution adheres to a form of the ``cluster assumption", and 2) the semi-supervised learning paradigm can be regarded as a special case of our framework when there are no distributional shifts. We validate our claims through experiments conducted on a variety of synthetic and real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- Gradual Domain Adaptation via Manifold-Constrained Distributionally Robust OptimizationSeyed Amir Saberi, Amir Najafi, Amin Behjati, Ala Emrani 等NeurIPS 2024 · 被引用 3 次
- Complementary Benefits of Contrastive Learning and Self-Training Under Distribution ShiftSaurabh Garg, Amrith Setlur, Zachary C. Lipton, Sivaraman Balakrishnan 等NeurIPS 2023 · 被引用 13 次
- A Characterization of Semi-Supervised Adversarially Robust PAC LearnabilityIdan Attias, Steve Hanneke, Yishay MansourNeurIPS 2022 · 被引用 19 次
- Generalized Semi-Supervised Learning via Self-Supervised Feature AdaptationJiachen Liang, Ruibing Hou, Hong Chang, Bingpeng Ma 等NeurIPS 2023 · 被引用 7 次
- Distributionally Robust Classification for Multi-source Unsupervised Domain AdaptationSeonghwi Kim, Sungho Jo, Wooseok Ha, Minwoo ChaeICLR 2026 · 被引用 4 次
