Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
Rizhen Hu, Yuan Cao, Boao Kong, Mou Sun, Kun Yuan
摘要
Sparse Mixture-of-Experts (MoE) models scale Transformers efficiently but suffer from expert overlap-redundant representations across experts and routing ambiguity, resulting in severely underutilized model capacity. While architectural solutions like DeepSeekMoE promote specialization, they require substantial structural modifications and rely solely on intra-layer signals. In this paper, we propose two plug-and-play regularization losses that enhance MoE specialization and routing efficiency without modifying router or model architectures. First, an intralayer specialization loss penalizes cosine similarity between experts' SwiGLU activations on identical tokens, encouraging experts to specialize in complementary knowledge. Second, a crosslayer coupling loss maximizes joint Top-k routing probabilities across adjacent layers, establishing coherent expert pathways through network depth while reinforcing intra-layer expert specialization. Both losses are orthogonal to the standard load-balancing loss and compatible with both the shared-expert architecture in DeepSeekMoE and vanilla top-k MoE architectures. We implement both losses as a drop-in Megatron-LM module. Extensive experiments across pre-training, finetuning, and zero-shot benchmarks demonstrate consistent task gains, higher expert specialization, and lower-entropy routing; together, these improvements translate into faster inference via more stable expert pathways.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank StructureBoao Kong, Junzhu Liang, Yuxi Liu, Renjia Deng 等ICLR 2026 · 被引用 7 次
- Grouter: Decoupling Routing from Representation for Accelerated MoE TrainingYuqi Xu, Rizhen Hu, zihan liu, Mou Sun 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper20
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
相关 Paper
- ReMoE: Fully Differentiable Mixture-of-Experts with ReLU RoutingZiteng Wang, Jun Zhu, Jianfei ChenICLR 2025
- Advancing Expert Specialization for Better MoEHongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu 等NeurIPS 2025 · 被引用 48 次
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary LossAng Lv, Jin Ma, Yiyuan Ma, Siyuan QiaoICLR 2026 · 被引用 14 次
- Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEsLeyla Mirvakhabova, Babak Ehteshami Bejnordi, Gaurav Kumar, Hanxue Liang 等ICML 2026 · 被引用 1 次
- Soft Modality-Guided Expert Specialization in MoE-VLMsZi-Hao Bo, Yaqian Li, Anzhou Hou, Rinyoichi Takezoe 等CVPR 2026
