Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
Jan Ludziejewski, Maciej Pióro, Jakub Krajewski, Maciej Stefaniak, Michal Krutul, Jan Malasnicki, Marek Cygan, Piotr Sankowski, Kamil Adamczewski, Piotr Milos, Sebastian Jaszczur
Abstract
Mixture of Experts (MoE) architectures have significantly increased computational efficiency in both research and real-world applications of large-scale machine learning models. However, their scalability and efficiency under memory constraints remain relatively underexplored. In this work, we present joint scaling laws for dense and MoE models, incorporating key factors such as the number of active parameters, dataset size, and the number of experts. Our findings provide a principled framework for selecting the optimal MoE configuration under fixed memory and compute budgets. Surprisingly, we show that MoE models can be more memory-efficient than dense models, contradicting conventional wisdom. To derive and validate the theoretical predictions of our scaling laws, we conduct over 280 experiments with up to 2.7B active parameters and up to 5B total parameters. These results offer actionable insights for designing and deploying MoE models in practical large-scale training scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language ModelsChangxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu et al.ICLR 2026 · 45 citations
- Scaling with Collapse: Efficient and Predictable Training of LLM FamiliesShane Bergsma, Bin Claire Zhang, Nolan Simran Dey, Shaheer Muhammad et al.ICLR 2026 · 10 citations
- Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal ResourceHouyi Li, Ka Man Lo, Shijie Xuyang, Ziqi Wang et al.ICLR 2026 · 8 citations
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task PlanningYuan Pu, Yazhe Niu, Jia Tang, Junyu Xiong et al.ICLR 2026 · 6 citations
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech RecognitionUmberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen et al.NeurIPS 2025 · 5 citations
Builds on10
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningTed Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermis et al.ICLR 2024 · 169 citations
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal et al.NeurIPS 2024 · 168 citations
- Scaling Laws for Fine-Grained Mixture of ExpertsJan Ludziejewski, Jakub Krajewski, Kamil Adamczewski, Maciej Pióro et al.ICML 2024 · 149 citations
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 144 citations
Related papers
- Mixture of Parrots: Experts improve memorization more than reasoningSamy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu et al.ICLR 2025
- Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language ModelsSiqi Wang, Zhengyu Chen, Bei Li, Keqing He et al.EMNLP 2024 · 4 citations
- FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsJiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang et al.PPoPP 2022 · 97 citations
- Generalization and Scaling Laws for Mixture-of-ExpertsTransformersMansour ZOUBEIROU A MAYAKIICML 2026
- Scaling Laws for Upcycling Mixture-of-Experts Language ModelsSeng Pei Liew, Takuya Kato, Sho TakaseICML 2025
