Multilingual Routing in Mixture-of-Experts
Lucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu, Nanyun (Violet) Peng
摘要
Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In this work, we analyze expert routing patterns using parallel multilingual datasets and present highly interpretable layer-wise phenomena. We find that MoE models route tokens in language-specific ways in the early and late decoder layers but exhibit significant cross-lingual routing alignment in middle layers, mirroring parameter-sharing trends observed in dense LLMs. In particular, we reveal a clear, strong correlation between a model's performance in a given language and how similarly its tokens are routed to English in these layers. Extending beyond correlation, we explore inference-time interventions that induce higher cross-lingual routing alignment. We introduce a method that steers the router by promoting middle-layer task experts frequently activated in English, and it successfully increases multilingual performance. These 1-2% gains are remarkably consistent across two evaluation tasks, three models, and 15+ languages, especially given that these simple interventions override routers of extensively trained, state-of-the-art LLMs. In comparison, interventions outside of the middle layers or targeting multilingual-specialized experts only yield performance degradation. Altogether, we present numerous findings that explain how MoEs process non-English text and demonstrate that generalization is limited by the model’s ability to leverage language-universal experts in all languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt 等ICLR 2026 · 被引用 28 次
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 被引用 3 次
- Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-ExpertsHaolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper27
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsFuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni 等ICML 2024 · 被引用 183 次
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu 等ACL 2024 · 被引用 171 次
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani 等ACL 2025 · 被引用 144 次
相关 Paper
- Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-ExpertsXue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang 等ACL 2025 · 被引用 9 次
- Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMsWengang Li, Lingqi Zhang, Toshio Endo, Mohamed WahibICLR 2026
- Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family ExpertsGuorui Zheng, Xidong Wang, Juhao Liang, Nuo Chen 等ICLR 2025
- R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-ExpertsZhongyang Li, Ziyue Li, Tianyi ZhouICML 2025
- Your Mixture-of-Experts LLM Is Secretly an Embedding Model for FreeZiyue Li, Tianyi ZhouICLR 2025
