Mixture Proportion Estimation and PU Learning: A Modern Approach
Saurabh Garg, Yifan Wu, Alexander J. Smola, Sivaraman Balakrishnan, Zachary C. Lipton
Abstract
Given only positive examples and unlabeled examples (from both positive and negative classes), we might hope nevertheless to estimate an accurate positive-versus-negative classifier. Formally, this task is broken down into two subtasks: (i) Mixture Proportion Estimation (MPE) -- determining the fraction of positive examples in the unlabeled data; and (ii) PU-learning -- given such an estimate, learning the desired positive-versus-negative classifier. Unfortunately, classical methods for both problems break down in high-dimensional settings. Meanwhile, recently proposed heuristics lack theoretical coherence and depend precariously on hyperparameter tuning. In this paper, we propose two simple techniques: Best Bin Estimation (BBE) (for MPE); and Conditional Value Ignoring Risk (CVIR), a simple objective for PU-learning. Both methods dominate previous approaches empirically, and for BBE, we establish formal guarantees that hold whenever we can train a model to cleanly separate out a small subset of positive examples. Our final algorithm (TED), alternates between the two procedures, significantly improving both our mixture proportion estimator and classifier
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 330a8257-6b6f-4aca-a60c-c817e15da421Cited by top-tier papers25
- Domain Adaptation under Open Set Label ShiftSaurabh Garg, Sivaraman Balakrishnan, Zachary C. LiptonNeurIPS 2022 · 57 citations
- Beyond Myopia: Learning from Positive and Unlabeled Data through Holistic Predictive TrendsXinrui Wang, Wenhai Wan, Chuanxing Geng, Shaoyuan Li et al.NeurIPS 2023 · 24 citations
- Federated Learning with Positive and Unlabeled DataXinyang Lin, Hanting Chen, Yixing Xu, Chao Xu et al.ICML 2022 · 23 citations
- Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label ConfigurationsHao Chen, Ankit Shah, Jindong Wang, Ran Tao et al.NeurIPS 2024 · 22 citations
- A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of LabelingYe Wang, Xinxin Liu, Wenxin Hu, Tao ZhangEMNLP 2022 · 18 citations
Builds on1
Related papers
- Balancing Positive and Negative Classification Error Rates in Positive-Unlabeled LearningXiming Li, Yuanchao Dai, Bing Wang, Changchun Li et al.NeurIPS 2025 · 3 citations
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 53 citations
- Rethinking Class-Prior Estimation for Positive-Unlabeled LearningYu Yao, Tongliang Liu, Bo Han, Mingming Gong et al.ICLR 2022 · 24 citations
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao et al.NeurIPS 2020 · 76 citations
- Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified PerspectiveQianqiao Liang, Mengying Zhu, Yan Wang, Xiuyuan Wang et al.AAAI 2023 · 4 citations
