Sparse MoE with Language Guided Routing for Multilingual Machine Translation
Xinyu Zhao, Xuxi Chen, Yu Cheng, Tianlong Chen
Abstract
Sparse Mixture-of-Experts (SMoE) has gained increasing popularity as a promising framework for scaling up multilingual machine translation (MMT) models with negligible extra computational overhead. However, current SMoE solutions neglect the intrinsic structures of the MMT problem: (a) Linguistics Hierarchy. Languages are naturally grouped according to their linguistic properties such as language families, phonological features, etc; (b) Language Complexity. Learning difficulties vary for different languages due to their available resources, grammar complexity etc. Therefore, routing a fixed number of experts (e.g., 1 or 2 experts in usual) only at the word level leads to inferior performance. To fill in the missing puzzle, we propose Lingual-SMoE by equipping the SMoE with adaptive and linguistics-guided routing policies. Specifically, it (1) extracts language representations to incorporate linguistic knowledge and uses them to allocate experts into different groups; (2) determines the number of activated experts for each target language in an adaptive and automatic manner, according to their difficulty level determined by data abundance, which aims to mitigate the potential over-/under-fitting problems of learning easy/difficult translations. Sufficient experimental studies on MMT benchmarks with 16, 50, 100 languages and various network architectures, consistently validate the superior performance of our proposals. For instance, Lingual-SMoE outperforms its dense counterpart by over 5% BLEU scores on the OPUS-100 dataset. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5604bd42-6cc9-40fb-90cb-2a475c3842efCited by top-tier papers7
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma et al.NeurIPS 2024 · 42 citations
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu et al.ICLR 2026 · 34 citations
- SimMLM: A Simple Framework for Multi-Modal Learning with Missing ModalitySijie Li, Chen Chen, Jungong HanICCV 2025 · 14 citations
- SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMsMikołaj Zasada, Łukasz Struski, Jacek Tabor, Marcin KurdzielICML 2026 · 2 citations
- Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction TuningSugyeong Eo, Jung Jun Lee, Chanjun Park, Heuiseok LimEMNLP 2025
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksShangjie Li, Xiangpeng Wei, Shaolin Zhu, Jun Xie et al.EMNLP 2023 · 4 citations
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine TranslationYunlong Liang, Fandong Meng, Jie ZhouACL 2025 · 1 citation
- SCoMoE: Efficient Mixtures of Experts with Structured CommunicationZhiyuan Zeng, Deyi XiongICLR 2023
- How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-ExpertsSumin Park, Noseong ParkAAAI 2026
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li et al.ACM MM 2025 · 1 citation
