On Linear Mode Connectivity of Mixture-of-Experts Architectures
Viet-Hoang Tran, Van-Hoan Trinh, Khanh Vinh Bui, Tan M. Nguyen
Abstract
Linear Mode Connectivity (LMC) is a notable phenomenon in the loss landscapes of neural networks, wherein independently trained models have been observed to be connected-up to permutation symmetries-by linear paths in parameter space along which the loss remains consistently low. This observation challenges classical views of non-convex optimization and has implications for model ensembling, generalization, and our understanding of neural loss geometry. Inspired by recent studies on LMC in standard neural networks, we systematically investigate this phenomenon within Mixture-of-Experts (MoE) architectures-a class of models known for their scalability and computational efficiency, which combine traditional neural networks-referred to as experts-through a learnable gating mechanism. We begin by conducting a comprehensive analysis of both dense and sparse gating regimes, demonstrating that the symmetries inherent to MoE architectures are fully characterized by permutations acting on both the expert components and the gating function. Building on these foundational findings, we propose a matching algorithm that enables alignment between independently trained MoEs, thereby facilitating the discovery of LMC. Finally, we empirically validate the presence of LMC using our proposed algorithm across diverse MoE configurations-including dense, sparse, and shared-expert variants-under a wide range of model settings and datasets of varying scales and modalities. Our results confirm the existence of LMC in MoE architectures and offer fundamental insights into the functional landscape and optimization dynamics of deep learning models. The code is publicly available at https://github.com/MLResearchX/lmc-moe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 831681cc-a853-47f2-9d2d-adb3a9860e5cCited by top-tier papers7
- Expert Merging in Sparse Mixture of Experts with Nash BargainingDung Viet Nguyen, Anh Nguyen Thi, Minh Hoang Nguyen, Luc Nguyen et al.ICLR 2026 · 3 citations
- Tree-Sliced Entropy Partial TransportViet-Hoang Tran, Thanh Tran, Thanh T. Chu, Tam Le et al.NeurIPS 2025 · 3 citations
- Robustness of Mixtures of Experts to Feature NoiseDong Sun, Rahul Nittala, Rebekka BurkholzICML 2026 · 1 citation
- Quasi-Equivariant MetanetworksViet-Hoang Tran, An Nguyen The, Benoît Guérand, Thieu Vo et al.ICLR 2026 · 1 citation
- Tree-sliced Sobolev IPMViet-Hoang Tran, Thanh Q. Tran, Thanh T. Chu, Duy-Tung Pham et al.ICLR 2026
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- Generalized Linear Mode Connectivity for TransformersAlexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto et al.NeurIPS 2025 · 18 citations
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 32 citations
- SD-MoE: Spectral Decomposition for Effective Expert SpecializationRuijun Huang, Fang DONG(董方), Xin Zhang, Anrui Chen et al.ICML 2026 · 2 citations
- Linear Mode Connectivity between Multiple Models modulo Permutation SymmetriesAkira Ito, Masanori Yamada, Atsutoshi KumagaiICML 2025
