Leveraging BERT and TFIDF Features for Short Text Clustering via Alignment-Promoting Co-Training
Zetong Li, Qinliang Su, Shijing Si, Jianxing Yu
Abstract
BERT and TFIDF features excel in capturing rich semantics and important words, respectively. Since most existing clustering methods are solely based on the BERT model, they often fall short in utilizing keyword information, which, however, is very useful in clustering short texts. In this paper, we propose a CO-Training Clustering (COTC) framework to make use of the collective strengths of BERT and TFIDF features. Specifically, we develop two modules responsible for the clustering of BERT and TFIDF features, respectively. We use the deep representations and cluster assignments from the TFIDF module outputs to guide the learning of the BERT module, seeking to align them at both the representation and cluster levels. Reversely, we also use the BERT module outputs to train the TFIDF module, thus leading to the mutual promotion. We then show that the alternating co-training framework can be placed under a unified joint training objective, which allows the two modules to be connected tightly and the training signals to be propagated efficiently. Experiments on eight benchmark datasets show that our method outperforms current SOTA methods significantly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 687a4e1a-e493-4fdb-b0c7-d01ef416bf4aCited by top-tier papers2
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- Neural Topic Modeling via Contextual and Graph Information FusionJiyuan Liu, Jiaxing Yan, Chunjiang Zhu, Xingyu Liu et al.EMNLP 2025
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 106 citations
- Robust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text ClusteringXiaolin Zheng, Mengling Hu, Weiming Liu, Chaochao Chen et al.ACL 2023 · 12 citations
Related papers
- TC-DWA: Text Clustering with Dual Word-Level AugmentationBo Cheng, Ximing Li, Yi ChangAAAI 2023 · 6 citations
- CEIL: A General Classification-Enhanced Iterative Learning Framework for Text ClusteringMingjun Zhao, Mengzhen Wang, Yinglong Ma, Di Niu et al.WWW 2023 · 2 citations
- Toward Open-domain Slot Filling via Self-supervised Co-trainingAdib Mosharrof, Moghis Fereidouni, A. B. SiddiqueWWW 2023 · 4 citations
- Graph-augmented and Over-smoothing-resistant Contrastive Clustering for Short TextZijian Zheng, Tao Ai, Yonghe LuAAAI 2026
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch et al.EMNLP 2020 · 144 citations
