Keep It on a Leash: Controllable Pseudo-label Generation Towards Realistic Long-Tailed Semi-Supervised Learning
Yaxin Hou, Bo Han, Yuheng Jia, Hui Liu, Junhui Hou
摘要
Current long-tailed semi-supervised learning methods assume that labeled data exhibit a long-tailed distribution, and unlabeled data adhere to a typical predefined distribution (i.e., long-tailed, uniform, or inverse long-tailed). However, the distribution of the unlabeled data is generally unknown and may follow an arbitrary distribution. To tackle this challenge, we propose a Controllable Pseudo-label Generation (CPG) framework, expanding the labeled dataset with the progressively identified reliable pseudo-labels from the unlabeled dataset and training the model on the updated labeled dataset with a known distribution, making it unaffected by the unlabeled data distribution. Specifically, CPG operates through a controllable self-reinforcing optimization cycle: (i) at each training step, our dynamic controllable filtering mechanism selectively incorporates reliable pseudo-labels from the unlabeled dataset into the labeled dataset, ensuring that the updated labeled dataset follows a known distribution; (ii) we then construct a Bayes-optimal classifier using logit adjustment based on the updated labeled data distribution; (iii) this improved classifier subsequently helps identify more reliable pseudo-labels in the next training step. We further theoretically prove that this optimization cycle can significantly reduce the generalization error under some conditions. Additionally, we propose a class-aware adaptive augmentation module to further improve the representation of minority classes, and an auxiliary branch to maximize data utilization by leveraging all labeled and unlabeled samples. Comprehensive evaluations on various commonly used benchmark datasets show that CPG achieves consistent improvements, surpassing state-of-the-art methods by up to \textbf{15.97%} in accuracy. The code is available at https://github.com/yaxinhou/CPG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DiCaP: Distribution-Calibrated Pseudo-labeling for Semi-Supervised Multi-Label LearningBo Han, Zhuoming Li, Xiaoyu Wang, Yaxin Hou 等AAAI 2026
- Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-Supervised LearningYaxin Hou, Jun Ma, Hanyang Li, Bo Han 等ICML 2026
它引用的顶会 Paper34
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 被引用 630 次
- Long-tailed Recognition by Routing Diverse Distribution-Aware ExpertsXudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu 等ICLR 2021 · 被引用 481 次
相关 Paper
- Towards Realistic Long-Tailed Semi-Supervised Learning: Consistency is All You NeedTong Wei, Kai GanCVPR 2023
- Continuous Contrastive Learning for Long-Tailed Semi-Supervised RecognitionZi-Hao Zhou, Siyuan Fang, Zi-Jing Zhou, Tong Wei 等NeurIPS 2024 · 被引用 19 次
- Three Heads Are Better than One: Complementary Experts for Long-Tailed Semi-supervised LearningChengcheng Ma, Ismail Elezi, Jiankang Deng, Weiming Dong 等AAAI 2024 · 被引用 20 次
- Re-distributing Biased Pseudo Labels for Semi-supervised Semantic Segmentation: A Baseline InvestigationRuifei He, Jihan Yang, Xiaojuan QiICCV 2021 · 被引用 149 次
- SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised LearningChaoqun Du, Yizeng Han, Gao HuangICML 2024 · 被引用 17 次
