Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
Wei Wang, Dong-Dong Wu, Ming Li, Jingxiong Zhang, Gang Niu, Masashi Sugiyama
摘要
Positive-unlabeled (PU) learning is a weakly supervised binary classification problem, in which the goal is to learn a binary classifier from only positive and unlabeled data, without access to negative data. In recent years, many PU learning algorithms have been developed to improve model performance. However, experimental settings are highly inconsistent, making it difficult to identify which algorithm performs better. In this paper, we propose the first PU learning benchmark to systematically compare PU learning algorithms. During our implementation, we identify subtle yet critical factors that affect the realistic and fair evaluation of PU learning algorithms. On the one hand, many PU learning algorithms rely on a validation set that includes negative data for model selection. This is unrealistic in traditional PU learning settings, where no negative data are available. To handle this problem, we systematically investigate model selection criteria for PU learning. On the other hand, PU learning involves different problem settings and corresponding solution families, i.e., the one-sample and two-sample settings. However, existing evaluation protocols are heavily biased towards the one-sample setting and neglect the significant difference between them. We identify the internal label shift problem of unlabeled training data for the one-sample setting and propose a simple yet effective calibration approach to ensure fair comparisons within and across families. We hope our framework will provide an accessible, realistic, and fair environment for evaluating PU learning algorithms in the future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Self-PU: Self Boosted and Calibrated Positive-Unlabeled TrainingXuxi Chen, Wuyang Chen, Tianlong Chen, Ye Yuan 等ICML 2020 · 被引用 100 次
- Multiscale Positive-Unlabeled Detection of AI-Generated TextsYuchuan Tian, Hanting Chen, Xutao Wang, Zheyuan Bai 等ICLR 2024 · 被引用 84 次
- Mixture Proportion Estimation and PU Learning: A Modern ApproachSaurabh Garg, Yifan Wu, Alexander J. Smola, Sivaraman Balakrishnan 等NeurIPS 2021 · 被引用 79 次
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao 等NeurIPS 2020 · 被引用 76 次
相关 Paper
- PU-BENCH: A Unified Benchmark for Rigorous and Reproducible PU LearningQiuyi Chen, Haiyang Zhang, Leqi Zhang, Changchun Li 等ICLR 2026
- A Closer Look to Positive-Unlabeled Learning from Fine-grained Perspectives: An Empirical StudyYuanchao Dai, Zhengzhang Hou, Changchun Li, Yuanbo Xu 等NeurIPS 2025 · 被引用 2 次
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 被引用 53 次
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu 等AAAI 2022 · 被引用 20 次
- PULNS: Positive-Unlabeled Learning with Effective Negative Sample SelectorChuan Luo, Pu Zhao, Chen Chen, Bo Qiao 等AAAI 2021 · 被引用 48 次
