EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering
Zhibin Duan, Hao Zhang, Chaojie Wang, Zhengjue Wang, Bo Chen, Mingyuan Zhou
Abstract
Natural language processing often faces the problem of data diversity such as different domains, themes, styles and so on. Therefore, a single language model (LM) is insufficient to learn all knowledge from diverse samples. To solve this problem, we firstly propose an autoencoding topic model with mixture prior (mATM) to perform clustering for the data, where the clusters defined in semantic space describe the data diversity. Having obtained the clustering assignment for each sample, we develop the ensemble LM (En-sLM) with the technique of weight modulation. Specifically, EnsLM contains a backbone which is adjusted by a few modulated weights to fit for different sample clusters. As a result, the backbone learns the shared knowledge among all clusters while modulated weights extract the cluster-specific features. EnsLM can be trained jointly with mATM with flexible LM backbone. We evaluate the effectiveness of both mATM and EnsLM on different language understanding and generative tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Mitigating Inconsistencies in Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Jiantao Zhou, Tianyi LiuEMNLP 2022 · 29 citations
- LEMoE: Advanced Mixture of Experts Adaptor for Lifelong Model Editing of Large Language ModelsRenzhi Wang, Piji LiEMNLP 2024 · 3 citations
- Serial Lifelong Editing via Mixture of Knowledge ExpertsYuJu Cheng, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank WangACL 2025
Builds on9
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Adversarial and Domain-Aware BERT for Cross-Domain Sentiment AnalysisChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi et al.ACL 2020 · 165 citations
- GAN Memory with No ForgettingYulai Cong, Miaoyun Zhao, Jianqiao Li, Sijia Wang et al.NeurIPS 2020 · 156 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Friendly Topic Assistant for Transformer Based Abstractive SummarizationZhengjue Wang, Zhibin Duan, Hao Zhang, Chaojie Wang et al.EMNLP 2020 · 48 citations
Related papers
- Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource RegimesYishi Xu, Jianqiao Sun, Yudi Su, Xinyang Liu et al.NeurIPS 2023 · 9 citations
- Task-Adaptive Pretrained Language Models via Clustered-Importance SamplingDavid Grangier, Simin Fan, Skyler Seto, Pierre AblinICLR 2025
- Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka EmbeddingsHans William Alexander Hanley, Zakir DurumericACL 2025
- Rethinking LLM Ensembling from the Perspective of Mixture ModelsJiale Fu, Yuchu Jiang, PeiJun Wu, Chonghan Liu et al.ICML 2026
- SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online FeedbackBo Lv, Nayu Liu, Chen Tang, Xin Liu et al.NeurIPS 2025 · 7 citations
