Breaking ReLU Barrier: Generalized MoEfication for Dense Pretrained Models
Jaeseong Lee, Seung-won Hwang, Wonpyo Park, Mingi Ji
Abstract
As the scale of language models (LMs) continues to grow, there is a heightened interest in reducing the inference cost associated with these models. Mixture-of-Experts (MoEs) present an efficient alternative to dense models, while the existing methods to convert pretrained dense models to MoEs is limited to ReLU-based models with natural sparsity. This paper introduces G-MoEfication, applicable to arbitrary dense models, where ReLU-based activation sparsity assumptions no longer hold. For generalizations, we encounter the dilemma of needing to zero-out deactivated experts, while also avoiding excessive zeroing-out to retain dense activation information. We publicly release our code 1 and report results conducted with mBERT, SantaCoder-1.1B, Phi-2-2.7B, and Falcon-7B demonstrating the efficacy of our approach in general scenarios: from multitask to multilingual, from fine-tuning to zero-shot evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99a7c887-d9f4-46ae-ba69-5f712e8457d3Cited by top-tier papers3
- Analytical FFN-to-MoE Restructuring via Activation Pattern AnalysisZehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu et al.ACL 2026 · 6 citations
- Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation LinkingTuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao et al.ASPLOS 2025 · 1 citation
- Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive GenerationYing Li, Zefang Wang, Zhaode Wang, Zhiwen Chen et al.ICML 2026
Builds on8
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural NetworksMark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev et al.ICML 2020 · 163 citations
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language ModelsIman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C. del Mundo et al.ICLR 2024 · 109 citations
- DropNet: Reducing Neural Network Complexity via Iterative PruningChong Min John Tan, Mehul MotaniICML 2020 · 74 citations
Related papers
- Masks Can be Learned as an Alternative to ExpertsPeiyu Liu, Tianwen Wei, Bo Zhu, Xin Zhao et al.ACL 2025 · 1 citation
- Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts ConversionFilip Szatkowski, Bartosz Wójcik, Mikolaj Piórczynski, Simone ScardapaneNeurIPS 2024 · 19 citations
- STUN: Structured-Then-Unstructured Pruning for Scalable MoE PruningJaeseong Lee, Seung-won Hwang, Aurick Qiao, Daniel F. Campos et al.ACL 2025
- Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-DesignRuisi Cai, Yeonju Ro, Geon-Woo Kim, Peihao Wang et al.NeurIPS 2024 · 21 citations
- ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation PatternsZiyu Zhao, Tong Zhu, Xin Yu, Zhi Zhang et al.ICML 2026
