Do We Really Need Permutations? Impact of Model Width on Linear Mode Connectivity
Akira Ito, Masanori Yamada, Daiki Chijiwa, Atsutoshi Kumagai
摘要
Recently, Ainsworth et al. empirically demonstrated that, given two independently trained models, applying a parameter permutation that preserves the input–output behavior allows the two models to be connected by a low-loss linear path. When such a path exists, the models are said to achieve linear mode connectivity (LMC). Prior studies, including Ainsworth et al. (2023), have reported that achieving LMC requires not only an appropriate permutation search but also sufficiently wide models (e.g., a 32 width multiplier for ResNet-20). This is broadly believed to be because increasing the model width ensures a large enough space of candidate permutations, increasing the chance of finding one that yields LMC. In this work, we empirically demonstrate that, even without any permutations, simply widening the models is sufficient for achieving LMC when using a suitable softmax temperature calibration. We further explain why this phenomenon arises by analyzing intermediate layer outputs. Specifically, we introduce layerwise exponentially weighted connectivity (LEWC), which states that the output of each layer of the merged model can be represented as an exponentially weighted sum of the outputs of the corresponding layers of the original models. Consequently the merged model's output matches that of an ensemble of the original models, facilitating LMC. To the best of our knowledge, this work is the first to show that widening the model not only facilitates nonlinear mode connectivity, as suggested in prior research, but also significantly increases the possibility of achieving linear mode connectivity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos 等ICLR 2020 · 被引用 1,368 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
相关 Paper
- Linear Mode Connectivity between Multiple Models modulo Permutation SymmetriesAkira Ito, Masanori Yamada, Atsutoshi KumagaiICML 2025
- Analysis of Linear Mode Connectivity via Permutation-Based Weight Matching: With Insights into Other Permutation Search MethodsAkira Ito, Masanori Yamada, Atsutoshi KumagaiICLR 2025
- Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature ConnectivityZhanpeng Zhou, Yongyi Yang, Xiaojiang Yang, Junchi Yan 等NeurIPS 2023 · 被引用 56 次
- Linear Mode Connectivity in Differentiable Tree EnsemblesRyuichi Kanoh, Mahito SugiyamaICLR 2025
- On Linear Mode Connectivity of Mixture-of-Experts ArchitecturesViet-Hoang Tran, Van-Hoan Trinh, Khanh Vinh Bui, Tan M. NguyenNeurIPS 2025 · 被引用 9 次
