Uncertainty-Aware Routing for Principled Alignment with MoE Dynamics
Yilong Chen, Junyuan Shang, Yuchen Feng, Zhenyu Zhang, Naibin Gu, Ziqi Wang, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang
Abstract
Mixture-of-Experts (MoE) is a cornerstone for scaling LLMs, yet its training dynamics remain poorly understood, often leading to sub-optimal specialization. Moving beyond static routing, we present a systematic study of the MoE lifecycle using Helmholtz Free Energy and Router Entropy . We identify a universal Three-Stage Phase Transition —Exploration, Symmetry Breaking, and Stabilization—marked by an Energy “Climb” and Plateau . This reflects Frustrated Ex-ploration , caused by structural interference between specialization drives and uniformity constraints. To address this, we propose Uncertainty-Aware Routing (UAR) , which aligns routing with the model’s epistemic state via: (1) Evidence-Triggered Expansion , increasing active experts for high-energy tokens, and (2) Epistemic Masking , applying load-balancing only in high-uncertainty regimes to shield mature experts. Experiments confirm UAR reduces perplexity and improves expert distinctiveness, offering a principled path toward thermodynamically aligned computation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- Hierarchical Mixture of Experts with Two-Stage OptimizationGleb Molodtsov, Alexander Miasnikov, Aleksandr BeznosikovKDD 2026 · 2 citations
- When Model Merging Breaks Routing: Training-Free Calibration for MoECanbin Huang, Tianyuan Shi, Xiaojun Quan, Jingang Wang et al.ICML 2026
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu et al.ICLR 2026 · 34 citations
- Phase-Aware Mixture of Experts for Agentic Reinforcement LearningYang Shengtian, Ziteng Cui, Shuo He, Yewen Li et al.ICML 2026 · 2 citations
- On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language ModelsChongyang Zhao, Mingsong Li, Haodong Lu, Dong GongCVPR 2026 · 3 citations
