Rethinking Class-Prior Estimation for Positive-Unlabeled Learning
Yu Yao, Tongliang Liu, Bo Han, Mingming Gong, Gang Niu, Masashi Sugiyama, Dacheng Tao
摘要
Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distributional-assumption-free CPE methods rely on a critical assumption that the support of the positive data distribution cannot be contained in the support of the negative data distribution. If this is violated, those CPE methods will systematically overestimate the class prior; it is even worse that we cannot verify the assumption based on the data. In this paper, we rethink CPE for PU learning-can we remove the assumption to make CPE always valid? We show an affirmative answer by proposing Regrouping CPE (ReCPE) that builds an auxiliary probability distribution such that the support of the positive data distribution is never contained in the support of the negative data distribution. ReCPE can work with any CPE method by treating it as the base method. Theoretically, ReCPE does not affect its base if the assumption already holds for the original probability distribution; otherwise, it reduces the positive bias of its base. Empirically, ReCPE improves all state-of-the-art CPE methods on various datasets, implying that the assumption has indeed been violated here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Binary Classification with Confidence DifferenceWei Wang, Lei Feng, Yuchen Jiang, Gang Niu 等NeurIPS 2023 · 被引用 20 次
- Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More PracticalWei Wang, Takashi Ishida, Yu-Jie Zhang, Gang Niu 等ICML 2024 · 被引用 12 次
- Mixture Proportion Estimation Beyond IrreducibilityYilun Zhu, Aaron Fjeldsted, Darren Holland, George Landon 等ICML 2023 · 被引用 12 次
- AUC Maximization under Positive Distribution ShiftAtsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama 等NeurIPS 2024 · 被引用 7 次
- Unraveling the Impact of Heterophilic Structures on Graph Positive-Unlabeled LearningYuhao Wu, Jiangchao Yao, Bo Han, Lina Yao 等ICML 2024 · 被引用 5 次
它引用的顶会 Paper5
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang 等NeurIPS 2020 · 被引用 329 次
- Understanding and Improving Early Stopping for Learning with Noisy LabelsYingbin Bai, Erkun Yang, Bo Han, Yanhua Yang 等NeurIPS 2021 · 被引用 307 次
- Dual T: Reducing Estimation Error for Transition Matrix in Label-noise LearningYu Yao, Tongliang Liu, Bo Han, Mingming Gong 等NeurIPS 2020 · 被引用 297 次
- Sample Selection with Uncertainty of Losses for Learning with Noisy LabelsXiaobo Xia, Tongliang Liu, Bo Han, Mingming Gong 等ICLR 2022 · 被引用 139 次
- Instance-dependent Label-noise Learning under a Structural Causal ModelYu Yao, Tongliang Liu, Mingming Gong, Bo Han 等NeurIPS 2021 · 被引用 100 次
相关 Paper
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 被引用 53 次
- Class Prior Estimation with Biased Positives and Unlabeled ExamplesShantanu Jain, Justin Delano, Himanshu Sharma, Predrag RadivojacAAAI 2020 · 被引用 15 次
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu 等AAAI 2022 · 被引用 20 次
- Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified PerspectiveQianqiao Liang, Mengying Zhu, Yan Wang, Xiuyuan Wang 等AAAI 2023 · 被引用 4 次
- Learning from positive and unlabeled examples -Finite size sample boundsFarnam Mansouri, Shai Ben-DavidNeurIPS 2025 · 被引用 6 次
