Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
Houyi Li, Ka Man Lo, Shijie Xuyang, Ziqi Wang, Wenzhen Zheng, Haocheng Zhang, Zhao Li, Shuigeng Zhou, Xiangyu Zhang, Daxin Jiang
Abstract
Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense architectures under strictly equal resource constraints -that is, when the total parameter count, training compute, and data budget are identical? This question remains under-explored despite its significant practical value and potential. In this paper, we propose a novel perspective and methodological framework to study this question thoroughly. First, we comprehensively investigate the architecture of MoEs and achieve an optimal model design that maximizes the performance. Based on this, we subsequently find that an MoE model with activation rate in an optimal region is able to outperform its dense counterpart under the same total parameter, training compute and data resource. More importantly, this optimal region remains consistent across different model sizes. Although additional amount of data turns out to be a trade-off for enhanced performance, we show that this can be resolved via reusing data. We validate our findings through extensive experiments, training nearly 200 language models at 2B scale and over 50 at 7B scale, cumulatively processing 50 trillion tokens. All model checkpoints are publicly available † .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a453bfea-ea1d-4e3a-a395-eb1eee9fa8c6Builds on16
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
Related papers
- Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language ModelsSiqi Wang, Zhengyu Chen, Bei Li, Keqing He et al.EMNLP 2024 · 4 citations
- Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceBaihui Liu, Kaiyuan Tian, Wei Wang, Zhaoning Zhang et al.ACL 2026
- Joint MoE Scaling Laws: Mixture of Experts Can Be Memory EfficientJan Ludziejewski, Maciej Pióro, Jakub Krajewski, Maciej Stefaniak et al.ICML 2025
- Scaling Laws for Upcycling Mixture-of-Experts Language ModelsSeng Pei Liew, Takuya Kato, Sho TakaseICML 2025
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language ModelsChangxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu et al.ICLR 2026 · 45 citations
