Positive-Unlabeled Learning with Extreme Scarcity of Labeled Positives
Yuanchao Dai, Ximing Li, Wei Wang, Changchun Li, Gang Niu, Masashi Sugiyama
摘要
Positive-Unlabeled (PU) learning is a weakly-supervised paradigm that trains a binary classifier from labeled positive and unlabeled instances. In PU risk estimation, the empirical risk consists of an unlabeled term and a positive term. In this paper, we observe that when labeled positives are scarce, the risk deviation is dominated by the generalization bound of the positive term, which is composed of a complexity term governed by Rademacher complexity and a concentration term governed by the uniform range bound, leading to estimator instability. Based on this observation, we theoretically derive the sufficient sample threshold, defined as the smallest number of labeled positives required to achieve a target excess risk with high probability, and reveal its explicit dependence on both components. Inspired by this insight, we propose ScalePU, which incorporates variance regularization to induce a restricted sub-hypothesis space with reduced Rademacher complexity, and geometric regularization to encourage compact clustering of positive samples with a tighter effective range. Theoretical analysis demonstrates that both mechanisms effectively lower the threshold through improvements to different components of the bound. Experiments on eight benchmark datasets validate the effectiveness of ScalePU, with significant improvements under extreme label scarcity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
- AdaMatch: A Unified Approach to Semi-Supervised Learning and Domain AdaptationDavid Berthelot, Rebecca Roelofs, Kihyuk Sohn, Nicholas Carlini 等ICLR 2022 · 被引用 180 次
- FreeMatch: Self-adaptive Thresholding for Semi-supervised LearningYidong Wang, Hao Chen, Qiang Heng, Wenxin Hou 等ICLR 2023 · 被引用 139 次
- Self-PU: Self Boosted and Calibrated Positive-Unlabeled TrainingXuxi Chen, Wuyang Chen, Tianlong Chen, Ye Yuan 等ICML 2020 · 被引用 100 次
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao 等NeurIPS 2020 · 被引用 76 次
相关 Paper
- A Closer Look to Positive-Unlabeled Learning from Fine-grained Perspectives: An Empirical StudyYuanchao Dai, Zhengzhang Hou, Changchun Li, Yuanbo Xu 等NeurIPS 2025 · 被引用 2 次
- Balancing Positive and Negative Classification Error Rates in Positive-Unlabeled LearningXiming Li, Yuanchao Dai, Bing Wang, Changchun Li 等NeurIPS 2025 · 被引用 3 次
- Learning from positive and unlabeled examples -Finite size sample boundsFarnam Mansouri, Shai Ben-DavidNeurIPS 2025 · 被引用 6 次
- Dist-PU: Positive-Unlabeled Learning from a Label Distribution PerspectiveYunrui Zhao, Qianqian Xu, Yangbangyan Jiang, Peisong Wen 等CVPR 2022 · 被引用 47 次
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 被引用 53 次
