Semi-Supervised Text Classification with Balanced Deep Representation Distributions
Changchun Li, Ximing Li, Jihong Ouyang
摘要
Semi-Supervised Text Classification (SSTC) mainly works under the spirit of self-training. They initialize the deep classifier by training over labeled texts; and then alternatively predict unlabeled texts as their pseudo-labels and train the deep classifier over the mixture of labeled and pseudo-labeled texts. Naturally, their performance is largely affected by the accuracy of pseudo-labels for unlabeled texts. Unfortunately, they often suffer from low accuracy because of the margin bias problem caused by the large difference between representation distributions of labels in SSTC. To alleviate this problem, we apply the angular margin loss, and perform Gaussian linear transformation to achieve balanced label angle variances, i.e., the variance of label angles of texts within the same label. More accuracy of predicted pseudo-labels can be achieved by constraining all label angle variances balanced, where they are estimated over both labeled and pseudo-labeled texts during self-training loops. With this insight, we propose a novel SSTC method, namely Semi-Supervised Text Classification with Balanced Deep representation Distributions (S 2 TC-BDD). To evaluate S 2 TC-BDD, we compare it against the state-of-theart SSTC methods. Empirical results demonstrate the effectiveness of S 2 TC-BDD, especially when the labeled texts are scarce.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Don't Stop Pretraining? Make Prompt-based Fine-tuning Powerful LearnerZhengxiang Shi, Aldo LipaniNeurIPS 2023 · 被引用 36 次
- Learning with Partial Labels from Semi-supervised PerspectiveXiming Li, Yuanzhi Jiang, Changchun Li, Yiyuan Wang 等AAAI 2023 · 被引用 22 次
- JointMatch: A Unified Approach for Diverse and Collaborative Pseudo-Labeling to Semi-Supervised Text ClassificationHenry Peng Zou, Cornelia CarageaEMNLP 2023 · 被引用 13 次
- Positive and Unlabeled Learning with Controlled Probability Boundary FenceChangchun Li, Yuanchao Dai, Lei Feng, Ximing Li 等ICML 2024 · 被引用 8 次
- TC-DWA: Text Clustering with Dual Word-Level AugmentationBo Cheng, Ximing Li, Yi ChangAAAI 2023 · 被引用 6 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Deep Representation Learning on Long-Tailed Data: A Learnable Embedding Augmentation PerspectiveJialun Liu, Yifan Sun, Chuchu Han, Zhaopeng Dou 等CVPR 2020
相关 Paper
- Semi-supervised Multi-label Learning with Balanced Binary Angular Margin LossXiming Li, Silong Liang, Changchun Li, Pengfei Wang 等NeurIPS 2024 · 被引用 4 次
- Prototype-Guided Pseudo Labeling for Semi-Supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Lihong Wang 等ACL 2023 · 被引用 28 次
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
- Debiased Self-Training for Semi-Supervised LearningBaixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan 等NeurIPS 2022 · 被引用 162 次
- Self-Paced Pairwise Representation Learning for Semi-Supervised Text ClassificationJunfan Chen, Richong Zhang, Jiarui Wang, Chunming Hu 等WWW 2024 · 被引用 1 次
