DiCaP: Distribution-Calibrated Pseudo-labeling for Semi-Supervised Multi-Label Learning
Bo Han, Zhuoming Li, Xiaoyu Wang, Yaxin Hou, Hui Liu, Junhui Hou, Yuheng Jia
Abstract
Semi-supervised multi-label learning (SSMLL) aims to address the challenge of limited labeled data in multi-label learning (MLL) by leveraging unlabeled data to improve the model’s performance. While pseudo-labeling has become a dominant strategy in SSMLL, most existing methods assign equal weights to all pseudo-labels regardless of their quality, which can amplify the impact of noisy or uncertain predictions and degrade the overall performance. In this paper, we theoretically verify that the optimal weight for a pseudo-label should reflect its correctness likelihood. Empirically, we observe that on the same dataset, the correctness likelihood distribution of unlabeled data remains stable, even as the number of labeled training samples varies. Building on this insight, we propose Distribution-Calibrated Pseudo-labeling (DiCaP), a correctness-aware framework that estimates posterior precision to calibrate pseudo-label weights. We further introduce a dual-thresholding mechanism to separate confident and ambiguous regions: confident samples are pseudo-labeled and weighted accordingly, while ambiguous ones are explored by unsupervised contrastive learning. Experiments conducted on multiple benchmark datasets verify that our method achieves consistent improvements, surpassing state-of-the-art methods by up to 4.27%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30831d16-4a82-45bc-9584-42abeb7c7ebeBuilds on15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- Class-Distribution-Aware Pseudo-Labeling for Semi-Supervised Multi-Label LearningMing-Kun Xie, Jiahao Xiao, Hao-Zhe Liu, Gang Niu et al.NeurIPS 2023 · 55 citations
- Dual Relation Semi-Supervised Multi-Label LearningLichen Wang, Yunyu Liu, Can Qin, Gan Sun et al.AAAI 2020 · 47 citations
Related papers
- Correlation-Induced Label Prior for Semi-Supervised Multi-Label LearningBiao Liu, Ning Xu, Xiangyu Fang, Xin GengICML 2024
- Asymmetric Beta Loss for Evidence-Based Safe Semi-Supervised Multi-Label LearningHao-Zhe Liu, Ming-Kun Xie, Chen-Chen Zong, Sheng-Jun HuangKDD 2024 · 1 citation
- DAW: Exploring the Better Weighting Function for Semi-supervised Semantic SegmentationRui Sun, Huayu Mai, Tianzhu Zhang, Feng WuNeurIPS 2023 · 40 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- Rethinking Confidence Scores and Thresholds in Pseudolabeling-based SSLHarit Vishwakarma, Yi Chen, Satya Sai Srinath Namburi GNVV, Sui Jiet Tay et al.ICML 2025
