Rethinking Class-Prior Estimation for Positive-Unlabeled Learning
Yu Yao, Tongliang Liu, Bo Han, Mingming Gong, Gang Niu, Masashi Sugiyama, Dacheng Tao
Abstract
Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distributional-assumption-free CPE methods rely on a critical assumption that the support of the positive data distribution cannot be contained in the support of the negative data distribution. If this is violated, those CPE methods will systematically overestimate the class prior; it is even worse that we cannot verify the assumption based on the data. In this paper, we rethink CPE for PU learning-can we remove the assumption to make CPE always valid? We show an affirmative answer by proposing Regrouping CPE (ReCPE) that builds an auxiliary probability distribution such that the support of the positive data distribution is never contained in the support of the negative data distribution. ReCPE can work with any CPE method by treating it as the base method. Theoretically, ReCPE does not affect its base if the assumption already holds for the original probability distribution; otherwise, it reduces the positive bias of its base. Empirically, ReCPE improves all state-of-the-art CPE methods on various datasets, implying that the assumption has indeed been violated here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Binary Classification with Confidence DifferenceWei Wang, Lei Feng, Yuchen Jiang, Gang Niu et al.NeurIPS 2023 · 20 citations
- Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More PracticalWei Wang, Takashi Ishida, Yu-Jie Zhang, Gang Niu et al.ICML 2024 · 12 citations
- Mixture Proportion Estimation Beyond IrreducibilityYilun Zhu, Aaron Fjeldsted, Darren Holland, George Landon et al.ICML 2023 · 12 citations
- AUC Maximization under Positive Distribution ShiftAtsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama et al.NeurIPS 2024 · 7 citations
- Unraveling the Impact of Heterophilic Structures on Graph Positive-Unlabeled LearningYuhao Wu, Jiangchao Yao, Bo Han, Lina Yao et al.ICML 2024 · 5 citations
Builds on5
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang et al.NeurIPS 2020 · 329 citations
- Understanding and Improving Early Stopping for Learning with Noisy LabelsYingbin Bai, Erkun Yang, Bo Han, Yanhua Yang et al.NeurIPS 2021 · 307 citations
- Dual T: Reducing Estimation Error for Transition Matrix in Label-noise LearningYu Yao, Tongliang Liu, Bo Han, Mingming Gong et al.NeurIPS 2020 · 297 citations
- Sample Selection with Uncertainty of Losses for Learning with Noisy LabelsXiaobo Xia, Tongliang Liu, Bo Han, Mingming Gong et al.ICLR 2022 · 139 citations
- Instance-dependent Label-noise Learning under a Structural Causal ModelYu Yao, Tongliang Liu, Mingming Gong, Bo Han et al.NeurIPS 2021 · 100 citations
Related papers
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 53 citations
- Class Prior Estimation with Biased Positives and Unlabeled ExamplesShantanu Jain, Justin Delano, Himanshu Sharma, Predrag RadivojacAAAI 2020 · 15 citations
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu et al.AAAI 2022 · 20 citations
- Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified PerspectiveQianqiao Liang, Mengying Zhu, Yan Wang, Xiuyuan Wang et al.AAAI 2023 · 4 citations
- Learning from positive and unlabeled examples -Finite size sample boundsFarnam Mansouri, Shai Ben-DavidNeurIPS 2025 · 6 citations
