EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering
Zhibin Duan, Hao Zhang, Chaojie Wang, Zhengjue Wang, Bo Chen, Mingyuan Zhou
摘要
Natural language processing often faces the problem of data diversity such as different domains, themes, styles and so on. Therefore, a single language model (LM) is insufficient to learn all knowledge from diverse samples. To solve this problem, we firstly propose an autoencoding topic model with mixture prior (mATM) to perform clustering for the data, where the clusters defined in semantic space describe the data diversity. Having obtained the clustering assignment for each sample, we develop the ensemble LM (En-sLM) with the technique of weight modulation. Specifically, EnsLM contains a backbone which is adjusted by a few modulated weights to fit for different sample clusters. As a result, the backbone learns the shared knowledge among all clusters while modulated weights extract the cluster-specific features. EnsLM can be trained jointly with mATM with flexible LM backbone. We evaluate the effectiveness of both mATM and EnsLM on different language understanding and generative tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mitigating Inconsistencies in Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Jiantao Zhou, Tianyi LiuEMNLP 2022 · 被引用 29 次
- LEMoE: Advanced Mixture of Experts Adaptor for Lifelong Model Editing of Large Language ModelsRenzhi Wang, Piji LiEMNLP 2024 · 被引用 3 次
- Serial Lifelong Editing via Mixture of Knowledge ExpertsYuJu Cheng, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank WangACL 2025
它引用的顶会 Paper9
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 被引用 569 次
- Adversarial and Domain-Aware BERT for Cross-Domain Sentiment AnalysisChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi 等ACL 2020 · 被引用 165 次
- GAN Memory with No ForgettingYulai Cong, Miaoyun Zhao, Jianqiao Li, Sijia Wang 等NeurIPS 2020 · 被引用 156 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Friendly Topic Assistant for Transformer Based Abstractive SummarizationZhengjue Wang, Zhibin Duan, Hao Zhang, Chaojie Wang 等EMNLP 2020 · 被引用 48 次
相关 Paper
- Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource RegimesYishi Xu, Jianqiao Sun, Yudi Su, Xinyang Liu 等NeurIPS 2023 · 被引用 9 次
- Task-Adaptive Pretrained Language Models via Clustered-Importance SamplingDavid Grangier, Simin Fan, Skyler Seto, Pierre AblinICLR 2025
- Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka EmbeddingsHans William Alexander Hanley, Zakir DurumericACL 2025
- Rethinking LLM Ensembling from the Perspective of Mixture ModelsJiale Fu, Yuchu Jiang, PeiJun Wu, Chonghan Liu 等ICML 2026
- SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online FeedbackBo Lv, Nayu Liu, Chen Tang, Xin Liu 等NeurIPS 2025 · 被引用 7 次
