Contrast-Enhanced Semi-supervised Text Classification with Few Labels
Austin Cheng-Yun Tsai, Sheng-Ya Lin, Li-Chen Fu
摘要
Traditional text classification requires thousands of annotated data or an additional Neural Machine Translation (NMT) system, which are expensive to obtain in real applications. This paper presents a Contrast-Enhanced Semi-supervised Text Classification (CEST) framework under label-limited settings without incorporating any NMT systems. We propose a certainty-driven sample selection method and a contrast-enhanced similarity graph to utilize data more efficiently in self-training, alleviating the annotation-starving problem. The graph imposes a smoothness constraint on the unlabeled data to improve the coherence and the accuracy of pseudo-labels. Moreover, CEST formulates the training as a “learning from noisy labels” problem and performs the optimization accordingly. A salient feature of this formulation is the explicit suppression of the severe error propagation problem in conventional semi-supervised learning. With solely 30 labeled data per class for both training and validation dataset, CEST outperforms the previous state-of-the-art algorithms by 2.11% accuracy and only falls within the 3.04% accuracy range of fully-supervised pre-training language model fine-tuning on thousands of labeled data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Neighborhood-Regularized Self-Training for Learning with Few LabelsRan Xu, Yue Yu, Hejie Cui, Xuan Kan 等AAAI 2023 · 被引用 29 次
- Prototype-Guided Pseudo Labeling for Semi-Supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Lihong Wang 等ACL 2023 · 被引用 28 次
- Uncertainty-Aware Self-Training for Low-Resource Neural Sequence LabelingJianing Wang, Chengyu Wang, Jun Huang, Ming Gao 等AAAI 2023 · 被引用 5 次
- Open-Set Semi-Supervised Text Classification via Adversarial Disagreement MaximizationJunfan Chen, Richong Zhang, Junchi Chen, Chunming HuACL 2024 · 被引用 1 次
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
它引用的顶会 Paper6
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Understanding Self-Training for Gradual Domain AdaptationAnanya Kumar, Tengyu Ma, Percy LiangICML 2020 · 被引用 266 次
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong 等EMNLP 2020 · 被引用 203 次
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
相关 Paper
- Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language ModelsChristopher Schröder, Gerhard HeyerEMNLP 2024 · 被引用 2 次
- CEIL: A General Classification-Enhanced Iterative Learning Framework for Text ClusteringMingjun Zhao, Mengzhen Wang, Yinglong Ma, Di Niu 等WWW 2023 · 被引用 2 次
- PESCO: Prompt-enhanced Self Contrastive Learning for Zero-shot Text ClassificationYau-Shian Wang, Ta-Chung Chi, Ruohong Zhang, Yiming YangACL 2023 · 被引用 17 次
- Weakly-supervised Text Classification Based on Keyword GraphLu Zhang, Jiandong Ding, Yi Xu, Yingyao Liu 等EMNLP 2021 · 被引用 46 次
- CELDA: Leveraging Black-box Language Model as Enhanced Classifier without LabelsHyunsoo Cho, Youna Kim, Sang-goo LeeACL 2023 · 被引用 1 次
