Consistent Multi-Class Classification from Multiple Unlabeled Datasets
Zixi Wei, Senlin Shu, Yuzhou Cao, Hongxin Wei, Bo An, Lei Feng
Abstract
Weakly supervised learning aims to construct effective predictive models from imperfectly labeled data. The recent trend of weakly supervised learning has focused on how to learn an accurate classifier from completely unlabeled data, given little supervised information such as class priors. In this paper, we consider a newly proposed weakly supervised learning problem called multi-class classification from multiple unlabeled datasets, where only multiple sets of unlabeled data and their class priors (i.e., the proportion of each class) are provided for training the classifier. To solve this problem, we first propose a classifier-consistent method (CCM) based on a probability transition function. However, CCM cannot guarantee risk consistency and lacks of purified supervision information during training. Therefore, we further propose a risk-consistent method (RCM) that progressively purifies supervision information during training by importance weighting. We provide comprehensive theoretical analyses for our methods to demonstrate the statistical consistency. Experimental results on multiple benchmark datasets across various settings demonstrate the superiority of our proposed methods.
• We provide comprehensive theoretical analyses for our proposed methods CCM and RCM to demonstrate their theoretical guarantees.
• We conduct extensive experiments on benchmark datasets with various settings. Experimental results demonstrate that CCM works well but RCM consistently outperforms CCM.
In this section, we introduce necessary notations, related studies, and the problem setting of our work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9d29f5f-c811-4ba5-b961-e95472400a3fCited by top-tier papers1
Ask how each one uses itBuilds on5
- Provably Consistent Partial-Label LearningLei Feng, Jiaqi Lv, Bo Han, Miao Xu et al.NeurIPS 2020 · 188 citations
- PiCO: Contrastive Label Disambiguation for Partial Label LearningHaobo Wang, Ruixuan Xiao, Yixuan Li, Lei Feng et al.ICLR 2022 · 169 citations
- Do We Need Zero Training Loss After Achieving Zero Training Error?Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu et al.ICML 2020 · 155 citations
- Binary Classification from Multiple Unlabeled Datasets via Surrogate Set ClassificationNan Lu, Shida Lei, Gang Niu, Issei Sato et al.ICML 2021 · 17 citations
- Learning from Label Proportions: A Mutual Contamination FrameworkClayton Scott, Jianxin ZhangNeurIPS 2020 · 12 citations
Related papers
- AUC Optimization from Multiple Unlabeled DatasetsZheng Xie, Yu Liu, Ming LiAAAI 2024 · 2 citations
- Adversarial Multi Class Learning under Weak Supervision with Performance GuaranteesAlessio Mazzetto, Cyrus Cousins, Dylan Sam, Stephen H. Bach et al.ICML 2021 · 39 citations
- Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More PracticalWei Wang, Takashi Ishida, Yu-Jie Zhang, Gang Niu et al.ICML 2024 · 12 citations
- Can Class-Priors Help Single-Positive Multi-Label Learning?Biao Liu, Ning Xu, Jie Wang, Xin GengNeurIPS 2025
- A Universal Unbiased Method for Classification from Aggregate ObservationsZixi Wei, Lei Feng, Bo Han, Tongliang Liu et al.ICML 2023 · 7 citations
