On Linear Mode Connectivity of Mixture-of-Experts Architectures
Viet-Hoang Tran, Van-Hoan Trinh, Khanh Vinh Bui, Tan M. Nguyen
摘要
Linear Mode Connectivity (LMC) is a notable phenomenon in the loss landscapes of neural networks, wherein independently trained models have been observed to be connected-up to permutation symmetries-by linear paths in parameter space along which the loss remains consistently low. This observation challenges classical views of non-convex optimization and has implications for model ensembling, generalization, and our understanding of neural loss geometry. Inspired by recent studies on LMC in standard neural networks, we systematically investigate this phenomenon within Mixture-of-Experts (MoE) architectures-a class of models known for their scalability and computational efficiency, which combine traditional neural networks-referred to as experts-through a learnable gating mechanism. We begin by conducting a comprehensive analysis of both dense and sparse gating regimes, demonstrating that the symmetries inherent to MoE architectures are fully characterized by permutations acting on both the expert components and the gating function. Building on these foundational findings, we propose a matching algorithm that enables alignment between independently trained MoEs, thereby facilitating the discovery of LMC. Finally, we empirically validate the presence of LMC using our proposed algorithm across diverse MoE configurations-including dense, sparse, and shared-expert variants-under a wide range of model settings and datasets of varying scales and modalities. Our results confirm the existence of LMC in MoE architectures and offer fundamental insights into the functional landscape and optimization dynamics of deep learning models. The code is publicly available at https://github.com/MLResearchX/lmc-moe .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Expert Merging in Sparse Mixture of Experts with Nash BargainingDung Viet Nguyen, Anh Nguyen Thi, Minh Hoang Nguyen, Luc Nguyen 等ICLR 2026 · 被引用 3 次
- Tree-Sliced Entropy Partial TransportViet-Hoang Tran, Thanh Tran, Thanh T. Chu, Tam Le 等NeurIPS 2025 · 被引用 3 次
- Robustness of Mixtures of Experts to Feature NoiseDong Sun, Rahul Nittala, Rebekka BurkholzICML 2026 · 被引用 1 次
- Quasi-Equivariant MetanetworksViet-Hoang Tran, An Nguyen The, Benoît Guérand, Thieu Vo 等ICLR 2026 · 被引用 1 次
- Tree-sliced Sobolev IPMViet-Hoang Tran, Thanh Q. Tran, Thanh T. Chu, Duy-Tung Pham 等ICLR 2026
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
相关 Paper
- Generalized Linear Mode Connectivity for TransformersAlexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto 等NeurIPS 2025 · 被引用 18 次
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu 等NeurIPS 2022 · 被引用 199 次
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 被引用 32 次
- SD-MoE: Spectral Decomposition for Effective Expert SpecializationRuijun Huang, Fang DONG(董方), Xin Zhang, Anrui Chen 等ICML 2026 · 被引用 2 次
- Linear Mode Connectivity between Multiple Models modulo Permutation SymmetriesAkira Ito, Masanori Yamada, Atsutoshi KumagaiICML 2025
