Contrast-Enhanced Semi-supervised Text Classification with Few Labels
Austin Cheng-Yun Tsai, Sheng-Ya Lin, Li-Chen Fu
Abstract
Traditional text classification requires thousands of annotated data or an additional Neural Machine Translation (NMT) system, which are expensive to obtain in real applications. This paper presents a Contrast-Enhanced Semi-supervised Text Classification (CEST) framework under label-limited settings without incorporating any NMT systems. We propose a certainty-driven sample selection method and a contrast-enhanced similarity graph to utilize data more efficiently in self-training, alleviating the annotation-starving problem. The graph imposes a smoothness constraint on the unlabeled data to improve the coherence and the accuracy of pseudo-labels. Moreover, CEST formulates the training as a “learning from noisy labels” problem and performs the optimization accordingly. A salient feature of this formulation is the explicit suppression of the severe error propagation problem in conventional semi-supervised learning. With solely 30 labeled data per class for both training and validation dataset, CEST outperforms the previous state-of-the-art algorithms by 2.11% accuracy and only falls within the 3.04% accuracy range of fully-supervised pre-training language model fine-tuning on thousands of labeled data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff9bf219-653e-4ec7-9fe0-d40b13832fecCited by top-tier papers5
- Neighborhood-Regularized Self-Training for Learning with Few LabelsRan Xu, Yue Yu, Hejie Cui, Xuan Kan et al.AAAI 2023 · 29 citations
- Prototype-Guided Pseudo Labeling for Semi-Supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Lihong Wang et al.ACL 2023 · 28 citations
- Uncertainty-Aware Self-Training for Low-Resource Neural Sequence LabelingJianing Wang, Chengyu Wang, Jun Huang, Ming Gao et al.AAAI 2023 · 5 citations
- Open-Set Semi-Supervised Text Classification via Adversarial Disagreement MaximizationJunfan Chen, Richong Zhang, Junchi Chen, Chunming HuACL 2024 · 1 citation
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
Builds on6
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Understanding Self-Training for Gradual Domain AdaptationAnanya Kumar, Tengyu Ma, Percy LiangICML 2020 · 266 citations
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
Related papers
- Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language ModelsChristopher Schröder, Gerhard HeyerEMNLP 2024 · 2 citations
- CEIL: A General Classification-Enhanced Iterative Learning Framework for Text ClusteringMingjun Zhao, Mengzhen Wang, Yinglong Ma, Di Niu et al.WWW 2023 · 2 citations
- PESCO: Prompt-enhanced Self Contrastive Learning for Zero-shot Text ClassificationYau-Shian Wang, Ta-Chung Chi, Ruohong Zhang, Yiming YangACL 2023 · 17 citations
- Weakly-supervised Text Classification Based on Keyword GraphLu Zhang, Jiandong Ding, Yi Xu, Yingyao Liu et al.EMNLP 2021 · 46 citations
- CELDA: Leveraging Black-box Language Model as Enhanced Classifier without LabelsHyunsoo Cho, Youna Kim, Sang-goo LeeACL 2023 · 1 citation
