Prototype-Guided Pseudo Labeling for Semi-Supervised Text Classification
Weiyi Yang, Richong Zhang, Junfan Chen, Lihong Wang, Jaein Kim
摘要
Semi-supervised text classification (SSTC) aims at text classification with few labeled data and massive unlabeled data. Recent works achieve this task by pseudo-labeling methods, with the belief that the unlabeled and labeled data have identical data distribution, and assign the unlabeled data with pseudo-labels as additional supervision. However, existing pseudo-labeling methods usually suffer from ambiguous categorical boundary issues when training the pseudo-labeling phase, and simply select pseudo-labels without considering the unbalanced categorical distribution of the unlabeled data, making it difficult to generate reliable pseudo-labels for each category. We propose a novel semi-supervised framework, namely ProtoS2, with prototypical cluster separation (PCS) and prototypical-center data selection (CDS) technology to address the issue. Particularly, PCS exploits categorical prototypes to assimilate instance representations within the same category, thus emphasizing low-density separation for the pseudo-labeled data to alleviate ambiguous boundaries. Besides, CDS selects central pseudo-labeled data considering the categorical distribution, avoiding the model from biasing on dominant categories. Empirical studies and extensive analysis with four benchmarks demonstrate the effectiveness of the proposed model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Inference-Time Dynamic Modality Selection for Incomplete Multimodal ClassificationSiyi Du, Xinzhe Luo, Declan O'regan, Chen QinICLR 2026 · 被引用 4 次
- Weakly-Supervised Audio-Visual Video Parsing with Prototype-Based Pseudo-LabelingKranthi Kumar Rachavarapu, Kalyan Ramakrishnan, A. N. RajagopalanCVPR 2024 · 被引用 3 次
- DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningWeifeng Jiang, Qianren Mao, Chenghua Lin, Jianxin Li 等EMNLP 2023 · 被引用 2 次
- Semi-Supervised Multimodal Classification Through Learning from Modal and Strategic ComplementaritiesJunchi Chen, Richong Zhang, Junfan ChenAAAI 2025 · 被引用 1 次
- Open-Set Semi-Supervised Text Classification via Adversarial Disagreement MaximizationJunfan Chen, Richong Zhang, Junchi Chen, Chunming HuACL 2024 · 被引用 1 次
它引用的顶会 Paper8
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Prototypical Contrastive Learning of Unsupervised RepresentationsJunnan Li, Pan Zhou, Caiming Xiong, Steven C. H. HoiICLR 2021 · 被引用 484 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Self-Tuning for Data-Efficient Deep LearningXimei Wang, Jinghan Gao, Mingsheng Long, Jianmin WangICML 2021 · 被引用 79 次
相关 Paper
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
- CIDC: Cluster Identification-Guided Dual Correction for Robust Short Text ClusteringYuhua Zhao, Zhixin Han, Xuan Li, Peiyu Xu 等WWW 2026
- Semi-Supervised Text Classification with Balanced Deep Representation DistributionsChangchun Li, Ximing Li, Jihong OuyangACL 2021
- Semi-supervised Semantic Segmentation via Prototypical Contrastive LearningZenggui Chen, Zhouhui LianACM MM 2022 · 被引用 14 次
- Robust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text ClusteringXiaolin Zheng, Mengling Hu, Weiming Liu, Chaochao Chen 等ACL 2023 · 被引用 12 次
