Weakly-supervised Text Classification Based on Keyword Graph
Lu Zhang, Jiandong Ding, Yi Xu, Yingyao Liu, Shuigeng Zhou
摘要
Weakly-supervised text classification has received much attention in recent years for it can alleviate the heavy burden of annotating massive data. Among them, keyword-driven methods are the mainstream where user-provided keywords are exploited to generate pseudolabels for unlabeled texts. However, existing methods treat keywords independently, thus ignore the correlation among them, which should be useful if properly exploited. In this paper, we propose a novel framework called ClassKG to explore keyword-keyword correlation on keyword graph by GNN. Our framework is an iterative process. In each iteration, we first construct a keyword graph, so the task of assigning pseudo labels is transformed to annotating keyword subgraphs. To improve the annotation quality, we introduce a self-supervised task to pretrain a subgraph annotator, and then finetune it. With the pseudo labels generated by the subgraph annotator, we then train a text classifier to classify the unlabeled texts. Finally, we re-extract keywords from the classified texts. Extensive experiments on both long-text and short-text datasets show that our method substantially outperforms the existing ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal SupervisionYunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li 等WWW 2025 · 被引用 53 次
- Zero-Shot Text Classification with Self-TrainingAriel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz 等EMNLP 2022 · 被引用 48 次
- DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled SamplesYi Xu, Jiandong Ding, Lu Zhang, Shuigeng ZhouNeurIPS 2021 · 被引用 34 次
- PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble TrainingYunyi Zhang, Minhao Jiang, Yu Meng, Yu Zhang 等EMNLP 2023 · 被引用 16 次
- Beyond prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering RepresentationsYu Fei, Zhao Meng, Ping Nie, Roger Wattenhofer 等EMNLP 2022 · 被引用 13 次
它引用的顶会 Paper10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
- GCC: Graph Contrastive Coding for Graph Neural Network Pre-TrainingJiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang 等KDD 2020 · 被引用 755 次
相关 Paper
- FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationTingyu Xia, Yue Wang, Yuan Tian, Yi ChangEMNLP 2022 · 被引用 1 次
- META: Metadata-Empowered Weak Supervision for Text ClassificationDheeraj Mekala, Xinyang Zhang, Jingbo ShangEMNLP 2020 · 被引用 34 次
- Unsupervised Clustering using Pseudo-semi-supervised LearningDivam Gupta, Ramachandran Ramjee, Nipun Kwatra, Muthian SivathanuICLR 2020 · 被引用 21 次
- Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed GraphsHuanjing Zhao, Beining Yang, Yukuo Cen, Junyu Ren 等KDD 2024 · 被引用 12 次
- Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated DataDheeraj Mekala, Varun Gangal, Jingbo ShangEMNLP 2021 · 被引用 20 次
