Positive-Unlabeled Learning with Extreme Scarcity of Labeled Positives
Yuanchao Dai, Ximing Li, Wei Wang, Changchun Li, Gang Niu, Masashi Sugiyama
Abstract
Positive-Unlabeled (PU) learning is a weakly-supervised paradigm that trains a binary classifier from labeled positive and unlabeled instances. In PU risk estimation, the empirical risk consists of an unlabeled term and a positive term. In this paper, we observe that when labeled positives are scarce, the risk deviation is dominated by the generalization bound of the positive term, which is composed of a complexity term governed by Rademacher complexity and a concentration term governed by the uniform range bound, leading to estimator instability. Based on this observation, we theoretically derive the sufficient sample threshold, defined as the smallest number of labeled positives required to achieve a target excess risk with high probability, and reveal its explicit dependence on both components. Inspired by this insight, we propose ScalePU, which incorporates variance regularization to induce a restricted sub-hypothesis space with reduced Rademacher complexity, and geometric regularization to encourage compact clustering of positive samples with a tighter effective range. Theoretical analysis demonstrates that both mechanisms effectively lower the threshold through improvements to different components of the bound. Experiments on eight benchmark datasets validate the effectiveness of ScalePU, with significant improvements under extreme label scarcity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fbe99e8-3023-4aaa-82c4-1f0080b97958Builds on14
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- AdaMatch: A Unified Approach to Semi-Supervised Learning and Domain AdaptationDavid Berthelot, Rebecca Roelofs, Kihyuk Sohn, Nicholas Carlini et al.ICLR 2022 · 180 citations
- FreeMatch: Self-adaptive Thresholding for Semi-supervised LearningYidong Wang, Hao Chen, Qiang Heng, Wenxin Hou et al.ICLR 2023 · 139 citations
- Self-PU: Self Boosted and Calibrated Positive-Unlabeled TrainingXuxi Chen, Wuyang Chen, Tianlong Chen, Ye Yuan et al.ICML 2020 · 100 citations
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao et al.NeurIPS 2020 · 76 citations
Related papers
- A Closer Look to Positive-Unlabeled Learning from Fine-grained Perspectives: An Empirical StudyYuanchao Dai, Zhengzhang Hou, Changchun Li, Yuanbo Xu et al.NeurIPS 2025 · 2 citations
- Balancing Positive and Negative Classification Error Rates in Positive-Unlabeled LearningXiming Li, Yuanchao Dai, Bing Wang, Changchun Li et al.NeurIPS 2025 · 3 citations
- Learning from positive and unlabeled examples -Finite size sample boundsFarnam Mansouri, Shai Ben-DavidNeurIPS 2025 · 6 citations
- Dist-PU: Positive-Unlabeled Learning from a Label Distribution PerspectiveYunrui Zhao, Qianqian Xu, Yangbangyan Jiang, Peisong Wen et al.CVPR 2022 · 47 citations
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 53 citations
