Mixture Proportion Estimation and PU Learning: A Modern Approach
Saurabh Garg, Yifan Wu, Alexander J. Smola, Sivaraman Balakrishnan, Zachary C. Lipton
摘要
Given only positive examples and unlabeled examples (from both positive and negative classes), we might hope nevertheless to estimate an accurate positive-versus-negative classifier. Formally, this task is broken down into two subtasks: (i) Mixture Proportion Estimation (MPE) -- determining the fraction of positive examples in the unlabeled data; and (ii) PU-learning -- given such an estimate, learning the desired positive-versus-negative classifier. Unfortunately, classical methods for both problems break down in high-dimensional settings. Meanwhile, recently proposed heuristics lack theoretical coherence and depend precariously on hyperparameter tuning. In this paper, we propose two simple techniques: Best Bin Estimation (BBE) (for MPE); and Conditional Value Ignoring Risk (CVIR), a simple objective for PU-learning. Both methods dominate previous approaches empirically, and for BBE, we establish formal guarantees that hold whenever we can train a model to cleanly separate out a small subset of positive examples. Our final algorithm (TED), alternates between the two procedures, significantly improving both our mixture proportion estimator and classifier
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Domain Adaptation under Open Set Label ShiftSaurabh Garg, Sivaraman Balakrishnan, Zachary C. LiptonNeurIPS 2022 · 被引用 57 次
- Beyond Myopia: Learning from Positive and Unlabeled Data through Holistic Predictive TrendsXinrui Wang, Wenhai Wan, Chuanxing Geng, Shaoyuan Li 等NeurIPS 2023 · 被引用 24 次
- Federated Learning with Positive and Unlabeled DataXinyang Lin, Hanting Chen, Yixing Xu, Chao Xu 等ICML 2022 · 被引用 23 次
- Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label ConfigurationsHao Chen, Ankit Shah, Jindong Wang, Ran Tao 等NeurIPS 2024 · 被引用 22 次
- A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of LabelingYe Wang, Xinxin Liu, Wenxin Hu, Tao ZhangEMNLP 2022 · 被引用 18 次
它引用的顶会 Paper1
相关 Paper
- Balancing Positive and Negative Classification Error Rates in Positive-Unlabeled LearningXiming Li, Yuanchao Dai, Bing Wang, Changchun Li 等NeurIPS 2025 · 被引用 3 次
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 被引用 53 次
- Rethinking Class-Prior Estimation for Positive-Unlabeled LearningYu Yao, Tongliang Liu, Bo Han, Mingming Gong 等ICLR 2022 · 被引用 24 次
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao 等NeurIPS 2020 · 被引用 76 次
- Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified PerspectiveQianqiao Liang, Mengying Zhu, Yan Wang, Xiuyuan Wang 等AAAI 2023 · 被引用 4 次
