SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
Chaoqun Du, Yizeng Han, Gao Huang
摘要
Recent advancements in semi-supervised learning have focused on a more realistic yet challenging task: addressing imbalances in labeled data while the class distribution of unlabeled data remains both unknown and potentially mismatched. Current approaches in this sphere often presuppose rigid assumptions regarding the class distribution of unlabeled data, thereby limiting the adaptability of models to only certain distribution ranges. In this study, we propose a novel approach, introducing a highly adaptable framework, designated as SimPro, which does not rely on any predefined assumptions about the distribution of unlabeled data. Our framework, grounded in a probabilistic model, innovatively refines the expectation-maximization (EM) algorithm by explicitly decoupling the modeling of conditional and marginal class distributions. This separation facilitates a closed-form solution for class distribution estimation during the maximization phase, leading to the formulation of a Bayes classifier. The Bayes classifier, in turn, enhances the quality of pseudo-labels in the expectation phase. Remarkably, the SimPro framework not only comes with theoretical guarantees but also is straightforward to implement. Moreover, we introduce two novel class distributions broadening the scope of the evaluation. Our method showcases consistent state-of-the-art performance across diverse benchmarks and data distribution scenarios. Our code is available at https://github.com/ LeapLabTHU/SimPro .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Improved Balanced Classification with Theoretically Grounded Loss FunctionsCorinna Cortes, Mehryar Mohri, Yutao ZhongNeurIPS 2025 · 被引用 19 次
- Keep It on a Leash: Controllable Pseudo-label Generation Towards Realistic Long-Tailed Semi-Supervised LearningYaxin Hou, Bo Han, Yuheng Jia, Hui Liu 等NeurIPS 2025 · 被引用 4 次
- Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and RelabelingXiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou 等AAAI 2026 · 被引用 1 次
- Sampling Control for Imbalanced Calibration in Semi-Supervised LearningSenmao Tian, Xiang Wei, Shunli ZhangAAAI 2026
- CoLA: Co-Calibrated Logit Adjustment for Long-Tailed Semi-Supervised LearningQian Shao, Qiyuan Chen, Jiahe Chen, Zepeng Li 等ICLR 2026
它引用的顶会 Paper11
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma 等NeurIPS 2020 · 被引用 861 次
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised LearningJaehyung Kim, Youngbum Hur, Sejun Park, Eunho Yang 等NeurIPS 2020 · 被引用 209 次
- ABC: Auxiliary Balanced Classifier for Class-imbalanced Semi-supervised LearningHyuck Lee, Seungjae Shin, Heeyoung KimNeurIPS 2021 · 被引用 131 次
相关 Paper
- Towards Realistic Long-Tailed Semi-Supervised Learning: Consistency is All You NeedTong Wei, Kai GanCVPR 2023
- DASO: Distribution-Aware Semantics-Oriented Pseudo-label for Imbalanced Semi-Supervised LearningYoungtaek Oh, Dong-Jin Kim, In So KweonCVPR 2022 · 被引用 80 次
- Class-Imbalanced Semi-Supervised Learning with Adaptive ThresholdingLan-Zhe Guo, Yufeng LiICML 2022 · 被引用 148 次
- Smoothed Adaptive Weighting for Imbalanced Semi-Supervised Learning: Improve Reliability Against Unknown Distribution DataZhengfeng Lai, Chao Wang, Henrry Gunawan, Sen-Ching S. Cheung 等ICML 2022 · 被引用 52 次
- Generalized Semi-Supervised Learning via Self-Supervised Feature AdaptationJiachen Liang, Ruibing Hou, Hong Chang, Bingpeng Ma 等NeurIPS 2023 · 被引用 7 次
