Prototype-Guided Pseudo Labeling for Semi-Supervised Text Classification
Weiyi Yang, Richong Zhang, Junfan Chen, Lihong Wang, Jaein Kim
Abstract
Semi-supervised text classification (SSTC) aims at text classification with few labeled data and massive unlabeled data. Recent works achieve this task by pseudo-labeling methods, with the belief that the unlabeled and labeled data have identical data distribution, and assign the unlabeled data with pseudo-labels as additional supervision. However, existing pseudo-labeling methods usually suffer from ambiguous categorical boundary issues when training the pseudo-labeling phase, and simply select pseudo-labels without considering the unbalanced categorical distribution of the unlabeled data, making it difficult to generate reliable pseudo-labels for each category. We propose a novel semi-supervised framework, namely ProtoS2, with prototypical cluster separation (PCS) and prototypical-center data selection (CDS) technology to address the issue. Particularly, PCS exploits categorical prototypes to assimilate instance representations within the same category, thus emphasizing low-density separation for the pseudo-labeled data to alleviate ambiguous boundaries. Besides, CDS selects central pseudo-labeled data considering the categorical distribution, avoiding the model from biasing on dominant categories. Empirical studies and extensive analysis with four benchmarks demonstrate the effectiveness of the proposed model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Inference-Time Dynamic Modality Selection for Incomplete Multimodal ClassificationSiyi Du, Xinzhe Luo, Declan O'regan, Chen QinICLR 2026 · 4 citations
- Weakly-Supervised Audio-Visual Video Parsing with Prototype-Based Pseudo-LabelingKranthi Kumar Rachavarapu, Kalyan Ramakrishnan, A. N. RajagopalanCVPR 2024 · 3 citations
- DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningWeifeng Jiang, Qianren Mao, Chenghua Lin, Jianxin Li et al.EMNLP 2023 · 2 citations
- Semi-Supervised Multimodal Classification Through Learning from Modal and Strategic ComplementaritiesJunchi Chen, Richong Zhang, Junfan ChenAAAI 2025 · 1 citation
- Open-Set Semi-Supervised Text Classification via Adversarial Disagreement MaximizationJunfan Chen, Richong Zhang, Junchi Chen, Chunming HuACL 2024 · 1 citation
Builds on8
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Prototypical Contrastive Learning of Unsupervised RepresentationsJunnan Li, Pan Zhou, Caiming Xiong, Steven C. H. HoiICLR 2021 · 484 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Self-Tuning for Data-Efficient Deep LearningXimei Wang, Jinghan Gao, Mingsheng Long, Jianmin WangICML 2021 · 79 citations
Related papers
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
- CIDC: Cluster Identification-Guided Dual Correction for Robust Short Text ClusteringYuhua Zhao, Zhixin Han, Xuan Li, Peiyu Xu et al.WWW 2026
- Semi-Supervised Text Classification with Balanced Deep Representation DistributionsChangchun Li, Ximing Li, Jihong OuyangACL 2021
- Semi-supervised Semantic Segmentation via Prototypical Contrastive LearningZenggui Chen, Zhouhui LianACM MM 2022 · 14 citations
- Robust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text ClusteringXiaolin Zheng, Mengling Hu, Weiming Liu, Chaochao Chen et al.ACL 2023 · 12 citations
