Text Classification Using Label Names Only: A Language Model Self-Training Approach
Yu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong, Heng Ji, Chao Zhang, Jiawei Han
摘要
Current text classification methods typically require a good number of human-labeled documents as training data, which can be costly and difficult to obtain in real applications. Humans can perform classification without seeing any labeled examples but only based on a small set of words describing the categories to be classified. In this paper, we explore the potential of only using the label name of each class to train classification models on unlabeled data, without using any labeled documents. We use pre-trained neural language models both as general linguistic knowledge sources for category understanding and as representation learning models for document classification. Our method (1) associates semantically related words with the label names, (2) finds category-indicative words and trains the model to predict their implied categories, and (3) generalizes the model via self-training. We show that our model achieves around 90% accuracy on four benchmark datasets including topic and sentiment classification without using any labeled documents but learning from unlabeled data supervised by at most 3 words (1 in most cases) per class as the label name 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary 等NeurIPS 2021 · 被引用 231 次
- Topic Discovery via Latent Space Clustering of Pretrained Language Model RepresentationsYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang 等WWW 2022 · 被引用 73 次
- Decoupling Knowledge from Memorization: Retrieval-augmented Prompt LearningXiang Chen, Lei Li, Ningyu Zhang, Xiaozhuan Liang 等NeurIPS 2022 · 被引用 68 次
- Weakly-Supervised Aspect-Based Sentiment Analysis via Joint Aspect-Sentiment Topic EmbeddingJiaxin Huang, Yu Meng, Fang Guo, Heng Ji 等EMNLP 2020 · 被引用 54 次
- TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal SupervisionYunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li 等WWW 2025 · 被引用 53 次
它引用的顶会 Paper7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 被引用 121 次
- Discriminative Topic Mining via Category-Name Guided Text EmbeddingYu Meng, Jiaxin Huang, Guangyuan Wang, Zihan Wang 等WWW 2020 · 被引用 80 次
- Hierarchical Topic Mining via Joint Spherical Tree and Text EmbeddingYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang 等KDD 2020 · 被引用 56 次
相关 Paper
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 被引用 1 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang 等AAAI 2024 · 被引用 7 次
- Beyond prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering RepresentationsYu Fei, Zhao Meng, Ping Nie, Roger Wattenhofer 等EMNLP 2022 · 被引用 13 次
- Zero-Shot Text Classification with Self-TrainingAriel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz 等EMNLP 2022 · 被引用 48 次
