TC-DWA: Text Clustering with Dual Word-Level Augmentation
Bo Cheng, Ximing Li, Yi Chang
Abstract
The pre-trained language models, e.g., ELMo and BERT, have recently achieved promising performance improvement in a wide range of NLP tasks, because they can output strong contextualized embedded features of words. Inspired by their great success, in this paper we aim to fine-tune them to effectively handle the text clustering task, i.e., a classic and fundamental challenge in machine learning. Accordingly, we propose a novel BERT-based method, namely Text Clustering with Dual Word-level Augmentation (TC-DWA). To be specific, we formulate a self-training objective and enhance it with a dual word-level augmentation technique. First, we suppose that each text contains several most informative words, called anchor words, supporting the full text semantics. We use the embedded features of anchor words as augmented features, which are selected by ranking the norm-based attention weights of words. Second, we formulate an expectation form of word augmentation, which is equivalent to generating infinite augmented features, and further suggest a tractable approximation of Taylor expansion for efficient optimization. To evaluate the effectiveness of TC-DWA, we conduct extensive experiments on several benchmark text datasets. The results demonstrate that TC-DWA consistently outperforms the state-of-the-art baseline methods. Code available: https://github.com/BoCheng-96/TC-DWA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec66e805-2670-4bc9-800a-1e30bc702e3fBuilds on12
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Structural Deep Clustering NetworkDeyu Bo, Xiao Wang, Chuan Shi, Meiqi Zhu et al.WWW 2020 · 645 citations
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
Related papers
- Leveraging BERT and TFIDF Features for Short Text Clustering via Alignment-Promoting Co-TrainingZetong Li, Qinliang Su, Shijing Si, Jianxing YuEMNLP 2024 · 2 citations
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsFei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai et al.EMNLP 2021 · 18 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- Cluster & Tune: Boost Cold Start Performance in Text ClassificationEyal Shnarch, Ariel Gera, Alon Halfon, Lena Dankin et al.ACL 2022 · 27 citations
- CEIL: A General Classification-Enhanced Iterative Learning Framework for Text ClusteringMingjun Zhao, Mengzhen Wang, Yinglong Ma, Di Niu et al.WWW 2023 · 2 citations
