When Are Solutions Connected in Deep Networks?
Quynh Nguyen, Pierre Bréchet, Marco Mondelli
摘要
The question of how and why the phenomenon of mode connectivity occurs in training deep neural networks has gained remarkable attention in the research community. From a theoretical perspective, two possible explanations have been proposed: (i) the loss function has connected sublevel sets, and (ii) the solutions found by stochastic gradient descent are dropout stable. While these explanations provide insights into the phenomenon, their assumptions are not always satisfied in practice. In particular, the first approach requires the network to have one layer with order of N neurons (N being the number of training samples), while the second one requires the loss to be almost invariant after removing half of the neurons at each layer (up to some rescaling of the remaining ones). In this work, we improve both conditions by exploiting the quality of the features at every intermediate layer together with a milder over-parameterization requirement. More specifically, we show that: (i) under generic assumptions on the features of intermediate layers, it suffices that the last two hidden layers have order of √ N neurons, and (ii) if subsets of features at each layer are linearly separable, then almost no overparameterization is needed to show the connectivity. Our experiments confirm that the proposed condition ensures the connectivity of solutions found by stochastic gradient descent, even in settings where the previous requirements do not hold. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterizationSimone Bombari, Mohammad Hossein Amani, Marco MondelliNeurIPS 2022 · 被引用 45 次
- FedCal: Achieving Local and Global Calibration in Federated Learning via Aggregated Parameterized ScalerHongyi Peng, Han Yu, Xiaoli Tang, Xiaoxiao LiICML 2024 · 被引用 12 次
- Connecting Independently Trained Modes via Layer-Wise ConnectivityYongding Tian, Zaid Al-Ars, Maksim Kitsak, H Peter HofsteeICML 2026 · 被引用 1 次
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Understanding Mode Connectivity via Parameter Space SymmetryBo Zhao, Nima Dehmamy, Robin Walters, Rose YuICML 2025
它引用的顶会 Paper6
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 被引用 82 次
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
- Truth or backpropaganda? An empirical investigation of deep learning theoryMicah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller 等ICLR 2020 · 被引用 36 次
- Piecewise linear activations substantially shape the loss surfaces of neural networksFengxiang He, Bohan Wang, Dacheng TaoICLR 2020 · 被引用 33 次
- Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU NetworksArsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin GolestaniICLR 2020 · 被引用 12 次
相关 Paper
- Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature ConnectivityZhanpeng Zhou, Yongyi Yang, Xiaojiang Yang, Junchi Yan 等NeurIPS 2023 · 被引用 56 次
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 被引用 52 次
- Input Space Mode Connectivity in Deep Neural NetworksJakub Vrábel, Ori Shem-Ur, Yaron Oz, David KruegerICLR 2025
- Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural CollapseArthur Jacot, Peter Súkeník, Zihan Wang, Marco MondelliICLR 2025
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential EquationsJiayao Zhang, Hua Wang, Weijie J. SuNeurIPS 2021 · 被引用 9 次
