Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
Zehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu, Wulong Liu, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu
Abstract
Scaling large language models (LLMs) improves performance but significantly increases inference costs, with feed-forward networks (FFNs) consuming the majority of computational resources. While Mixture-of-Experts (MoE) architectures can reduce this cost through sparse activation, restructuring existing dense models into MoEs typically requires extensive retraining on hundreds of billions of tokens. We propose an analytical post-training framework that rapidly restructures FFNs into sparse MoE architectures using only a small calibration dataset. The method analyzes neuron activation patterns to partition neurons into always-active shared experts and conditionally activated routed experts, then constructs a router analytically from representative neuron statistics, enabling immediate deployment or optional lightweight fine-tuning. This approach applies both to dense models and recursively to existing MoE models for hierarchical sparsity. Experiments demonstrate up to speedup in compute-bound scenarios with only minutes of processing and 2k-sample fine-tuning, outperforming methods requiring orders of magnitude more resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive GenerationYing Li, Zefang Wang, Zhaode Wang, Zhiwen Chen et al.ICML 2026
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM InferenceSiheng Xiong, Joe Zou, Faramarz Fekri, Yae Jee ChoICML 2026
Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- SliceGPT: Compress Large Language Models by Deleting Rows and ColumnsSaleh Ashkboos, Maximilian L. Croci, Marcelo Gennari Do Nascimento, Torsten Hoefler et al.ICLR 2024 · 346 citations
- SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer BlocksJiwon Song, Kyungseok Oh, Taesu Kim, Hyungjun Kim et al.ICML 2024 · 86 citations
Related papers
- Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-DesignRuisi Cai, Yeonju Ro, Geon-Woo Kim, Peihao Wang et al.NeurIPS 2024 · 21 citations
- Masks Can be Learned as an Alternative to ExpertsPeiyu Liu, Tianwen Wei, Bo Zhu, Xin Zhao et al.ACL 2025 · 1 citation
- DOT-MoE: Differentiable Optimal Transport for MoEficationUdbhav Bamba, Arnav Chavan, Aryamaan Thakur, Steven Teig et al.ICML 2026
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language ModelsXudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou et al.ACL 2024 · 16 citations
- DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMsMinxuan Lv, Zhenpeng Su, Leiyu Pan, Yizhe Xiong et al.EMNLP 2025
