Layer-wise linear mode connectivity
Linara Adilova, Maksym Andriushchenko, Michael Kamp, Asja Fischer, Martin Jaggi
摘要
Averaging neural network parameters is an intuitive method for fusing the knowledge of two independent models. It is most prominently used in federated learning. If models are averaged at the end of training, this can only lead to a good performing model if the loss surface of interest is very particular, i.e., the loss in the midpoint between the two models needs to be sufficiently low. This is impossible to guarantee for the non-convex losses of state-of-the-art networks. For averaging models trained on vastly different datasets, it was proposed to average only the parameters of particular layers or combinations of layers, resulting in better performing models. To get a better understanding of the effect of layer-wise averaging, we analyse the performance of the models that result from averaging single layers, or groups of layers. Based on our empirical and theoretical investigation, we introduce a novel notion of the layer-wise linear connectivity, and show that deep networks do not have layer-wise barriers between them. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- On the Emergence of Cross-Task Linearity in Pretraining-Finetuning ParadigmZhanpeng Zhou, Zijun Chen, Yilan Chen, Bo Zhang 等ICML 2024 · 被引用 26 次
- Generalized Linear Mode Connectivity for TransformersAlexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto 等NeurIPS 2025 · 被引用 18 次
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingTing Han, Linara Adilova, Henning Petzka, Jens Kleesiek 等NeurIPS 2025 · 被引用 9 次
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang 等NeurIPS 2025 · 被引用 7 次
- Flexible Sharpness-Aware Personalized Federated LearningXinda Xing, Qiugang Zhan, Xiurui Xie, Yuning Yang 等AAAI 2025 · 被引用 5 次
它引用的顶会 Paper32
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos 等ICLR 2020 · 被引用 1,368 次
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp 等ICLR 2021 · 被引用 1,166 次
相关 Paper
- Layer-Wise Adaptive Model Aggregation for Scalable Federated LearningSunwoo Lee, Tuo Zhang, Amir Salman AvestimehrAAAI 2023 · 被引用 87 次
- FedAvg Converges to Zero Training Loss Linearly for Overparameterized Multi-Layer Neural NetworksBingqing Song, Prashant Khanduri, Xinwei Zhang, Jinfeng Yi 等ICML 2023 · 被引用 10 次
- Widening the Network Mitigates the Impact of Data Heterogeneity on FedAvgLike Jian, Dong LiuICML 2025
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 被引用 330 次
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 被引用 32 次
