DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
Yongqi Huang, Peng Ye, Chenyu Huang, Jianjian Cao, Lin Zhang, Baopu Li, Gang Yu, Tao Chen
Abstract
Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Upcycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Advancing Expert Specialization for Better MoEHongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu et al.NeurIPS 2025 · 48 citations
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech RecognitionUmberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen et al.NeurIPS 2025 · 5 citations
- Enhancing Mixture-of-Experts Specialization via Cluster-Aware UpcyclingSanghyeok Chu, Pyunghwan Ahn, Gwangmo Song, Seung Hwan Kim et al.CVPR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-ExpertsQi Wang, Hanyang Peng, Yue YuAAAI 2026 · 1 citation
- Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initializationTaishi Nakamura, Takuya Akiba, Kazuki Fujii, Yusuke Oda et al.ICLR 2025
- BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsQizhen (Irene) Zhang, Nikolas Gritsch, Dwaraknath Gnaneshwar, Simon Guo et al.NeurIPS 2024 · 18 citations
- Scaling Laws for Upcycling Mixture-of-Experts Language ModelsSeng Pei Liew, Takuya Kato, Sho TakaseICML 2025
- Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-DesignRuisi Cai, Yeonju Ro, Geon-Woo Kim, Peihao Wang et al.NeurIPS 2024 · 21 citations
