Sparse MoE with Language Guided Routing for Multilingual Machine Translation
Xinyu Zhao, Xuxi Chen, Yu Cheng, Tianlong Chen
摘要
Sparse Mixture-of-Experts (SMoE) has gained increasing popularity as a promising framework for scaling up multilingual machine translation (MMT) models with negligible extra computational overhead. However, current SMoE solutions neglect the intrinsic structures of the MMT problem: (a) Linguistics Hierarchy. Languages are naturally grouped according to their linguistic properties such as language families, phonological features, etc; (b) Language Complexity. Learning difficulties vary for different languages due to their available resources, grammar complexity etc. Therefore, routing a fixed number of experts (e.g., 1 or 2 experts in usual) only at the word level leads to inferior performance. To fill in the missing puzzle, we propose Lingual-SMoE by equipping the SMoE with adaptive and linguistics-guided routing policies. Specifically, it (1) extracts language representations to incorporate linguistic knowledge and uses them to allocate experts into different groups; (2) determines the number of activated experts for each target language in an adaptive and automatic manner, according to their difficulty level determined by data abundance, which aims to mitigate the potential over-/under-fitting problems of learning easy/difficult translations. Sufficient experimental studies on MMT benchmarks with 16, 50, 100 languages and various network architectures, consistently validate the superior performance of our proposals. For instance, Lingual-SMoE outperforms its dense counterpart by over 5% BLEU scores on the OPUS-100 dataset. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma 等NeurIPS 2024 · 被引用 42 次
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu 等ICLR 2026 · 被引用 34 次
- SimMLM: A Simple Framework for Multi-Modal Learning with Missing ModalitySijie Li, Chen Chen, Jungong HanICCV 2025 · 被引用 14 次
- SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMsMikołaj Zasada, Łukasz Struski, Jacek Tabor, Marcin KurdzielICML 2026 · 被引用 2 次
- Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction TuningSugyeong Eo, Jung Jun Lee, Chanjun Park, Heuiseok LimEMNLP 2025
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksShangjie Li, Xiangpeng Wei, Shaolin Zhu, Jun Xie 等EMNLP 2023 · 被引用 4 次
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine TranslationYunlong Liang, Fandong Meng, Jie ZhouACL 2025 · 被引用 1 次
- SCoMoE: Efficient Mixtures of Experts with Structured CommunicationZhiyuan Zeng, Deyi XiongICLR 2023
- How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-ExpertsSumin Park, Noseong ParkAAAI 2026
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li 等ACM MM 2025 · 被引用 1 次
