Multilingual Routing in Mixture-of-Experts
Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu, Nanyun (Violet) Peng
Abstract
Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In this work, we analyze expert routing patterns using parallel multilingual datasets and present highly interpretable layer-wise phenomena. We find that MoE models route tokens in language-specific ways in the early and late decoder layers but exhibit significant cross-lingual routing alignment in middle layers, mirroring parameter-sharing trends observed in dense LLMs. In particular, we reveal a clear, strong correlation between a model's performance in a given language and how similarly its tokens are routed to English in these layers. Extending beyond correlation, we explore inference-time interventions that induce higher cross-lingual routing alignment. We introduce a method that steers the router by promoting middle-layer task experts frequently activated in English, and it successfully increases multilingual performance. These 1-2% gains are remarkably consistent across two evaluation tasks, three models, and 15+ languages, especially given that these simple interventions override routers of extensively trained, state-of-the-art LLMs. In comparison, interventions outside of the middle layers or targeting multilingual-specialized experts only yield performance degradation. Altogether, we present numerous findings that explain how MoEs process non-English text and demonstrate that generalization is limited by the model’s ability to leverage language-universal experts in all languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1896ae6d-54f0-4403-bc76-5e8d6b6ea616Cited by top-tier papers3
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt et al.ICLR 2026 · 28 citations
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 3 citations
- Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-ExpertsHaolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou et al.ACL 2026 · 3 citations
Builds on27
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsFuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni et al.ICML 2024 · 183 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
Related papers
- Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-ExpertsXue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang et al.ACL 2025 · 9 citations
- Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMsWengang Li, Lingqi Zhang, Toshio Endo, Mohamed WahibICLR 2026
- Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family ExpertsGuorui Zheng, Xidong Wang, Juhao Liang, Nuo Chen et al.ICLR 2025
- R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-ExpertsZhongyang Li, Ziyue Li, Tianyi ZhouICML 2025
- Your Mixture-of-Experts LLM Is Secretly an Embedding Model for FreeZiyue Li, Tianyi ZhouICLR 2025
