Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
Di Zhang, Xun Wu, Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong, Zewen Chi, Zhifang Sui, Furu Wei
Abstract
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for improving reasoning capabilities. However, training RLVR with Mixture-of-Experts (MoE) policies remains fragile and is often prone to reward collapse. We identify a MoE-specific source of instability, referred to as router shift (RS), where changes in expert routing across policy updates exacerbate off-policy mismatch. This effect leads to increasingly volatile importance-ratio signals and bursty clipping behavior, which consistently precede training collapse. Motivated by this diagnosis, we propose Router-Shift Policy Optimization (RSPO). RSPO computes a per-token router-shift ratio conditioned on the previously activated experts, applies stop-gradient and a lower-bound floor, and softly rescales importance ratios prior to clipping and aggregation. This design explicitly accounts for routing-induced distributional drift during off-policy optimization. We evaluate the effect of RSPO under two settings: a synthetic countdown task and real-world reasoning tasks on MATH and Code. Across both settings, RSPO achieves better performance and exhibits greater stability compared to recent MoE-based RLVR methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31f781b1-a812-4987-b9ec-e62ef270c3b8Cited by top-tier papers2
- Probing RLVR Training Instability through the Lens of Objective-Level HackingYiming Dong, Kun Fu, Haoyu Li, Xinyuan Zhu et al.ICML 2026
- PADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student LearningXinyue Peng, Yi Qian, Jiaojiao Lin, Wenjian Shao et al.ICML 2026
Builds on7
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 264 citations
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen et al.ICLR 2026 · 104 citations
Related papers
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference RoutersWenhan Ma, Hailin Zhang, Liang Zhao, Yifan Song et al.ICML 2026 · 51 citations
- Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary SignalsShuo Yang, Jinda Lu, Chiyu Ma, Kexin Huang et al.ICML 2026
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu et al.ICLR 2026 · 51 citations
- Balancing the Experts: Unlocking LoRA-MoE for GRPO via Mechanism-Aware RewardsChanglian Ma, Zizheng Huang, Xiangyu Zeng, Yi Wang et al.ICLR 2026
- MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language ModelsDohwan Ko, Jinyoung Park, Seoung Choi, Sanghyeok Lee et al.CVPR 2026 · 3 citations
