PU-BENCH: A Unified Benchmark for Rigorous and Reproducible PU Learning
Qiuyi Chen, Haiyang Zhang, Leqi Zhang, Changchun Li, Jia Wang, Wei Wang
Abstract
Positive-Unlabeled (PU) learning, a challenging paradigm for training binary classifiers from only positive and unlabeled samples, is fundamental to many applications. While numerous PU learning methods have been proposed, the research is systematically hindered by the lack of a standardized and comprehensive benchmark for rigorous evaluation. Inconsistent data generation, disparate experimental settings, and divergent metrics have led to irreproducible findings and unsubstantiated performance claims. To address this foundational challenge, we introduce PU-Bench, the first unified open-source benchmark for PU learning. PU-Bench provides: 1) a unified data generation pipeline to ensure consistent input across configurable sampling schemes, label ratios and labeling mechanisms; 2) an integrated framework of 18 state-of-the-art PU methods; and 3) standardized protocols for reproducible assessment. Through a large-scale empirical study on 8 diverse datasets (2880 evaluations in total), PU-Bench reveals a complex yet intuitive performance landscape, uncovering critical trade-offs between effectiveness and efficiency, and systematically mapping method robustness against variations in label frequency and selection bias. It is anticipated to serve as a foundational resource to catalyze reproducible, rigorous, and impactful research in the PU learning community. The source code is publicly available at https://github.com/XiXiphus/PU-Bench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5177bb72-51f0-4027-8446-327ff17552edCited by top-tier papers1
Ask how each one uses itBuilds on12
- Self-PU: Self Boosted and Calibrated Positive-Unlabeled TrainingXuxi Chen, Wuyang Chen, Tianlong Chen, Ye Yuan et al.ICML 2020 · 100 citations
- Mixture Proportion Estimation and PU Learning: A Modern ApproachSaurabh Garg, Yifan Wu, Alexander J. Smola, Sivaraman Balakrishnan et al.NeurIPS 2021 · 79 citations
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao et al.NeurIPS 2020 · 76 citations
- Predictive Adversarial Learning from Positive and Unlabeled DataWenpeng Hu, Ran Le, Bing Liu, Feng Ji et al.AAAI 2021 · 56 citations
- Dist-PU: Positive-Unlabeled Learning from a Label Distribution PerspectiveYunrui Zhao, Qianqian Xu, Yangbangyan Jiang, Peisong Wen et al.CVPR 2022 · 47 citations
Related papers
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu et al.AAAI 2022 · 20 citations
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 53 citations
- Learning from positive and unlabeled examples -Finite size sample boundsFarnam Mansouri, Shai Ben-DavidNeurIPS 2025 · 6 citations
- A Closer Look to Positive-Unlabeled Learning from Fine-grained Perspectives: An Empirical StudyYuanchao Dai, Zhengzhang Hou, Changchun Li, Yuanbo Xu et al.NeurIPS 2025 · 2 citations
- Positive-Unlabeled Learning by Latent Group-Aware Meta DisambiguationLin Long, Haobo Wang, Zhijie Jiang, Lei Feng et al.CVPR 2024 · 2 citations
