When Are Solutions Connected in Deep Networks?
Quynh Nguyen, Pierre Bréchet, Marco Mondelli
Abstract
The question of how and why the phenomenon of mode connectivity occurs in training deep neural networks has gained remarkable attention in the research community. From a theoretical perspective, two possible explanations have been proposed: (i) the loss function has connected sublevel sets, and (ii) the solutions found by stochastic gradient descent are dropout stable. While these explanations provide insights into the phenomenon, their assumptions are not always satisfied in practice. In particular, the first approach requires the network to have one layer with order of N neurons (N being the number of training samples), while the second one requires the loss to be almost invariant after removing half of the neurons at each layer (up to some rescaling of the remaining ones). In this work, we improve both conditions by exploiting the quality of the features at every intermediate layer together with a milder over-parameterization requirement. More specifically, we show that: (i) under generic assumptions on the features of intermediate layers, it suffices that the last two hidden layers have order of √ N neurons, and (ii) if subsets of features at each layer are linearly separable, then almost no overparameterization is needed to show the connectivity. Our experiments confirm that the proposed condition ensures the connectivity of solutions found by stochastic gradient descent, even in settings where the previous requirements do not hold. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c297beb2-ab86-4c3d-bbe5-9e884305028cCited by top-tier papers6
- Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterizationSimone Bombari, Mohammad Hossein Amani, Marco MondelliNeurIPS 2022 · 45 citations
- FedCal: Achieving Local and Global Calibration in Federated Learning via Aggregated Parameterized ScalerHongyi Peng, Han Yu, Xiaoli Tang, Xiaoxiao LiICML 2024 · 12 citations
- Connecting Independently Trained Modes via Layer-Wise ConnectivityYongding Tian, Zaid Al-Ars, Maksim Kitsak, H Peter HofsteeICML 2026 · 1 citation
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Understanding Mode Connectivity via Parameter Space SymmetryBo Zhao, Nima Dehmamy, Robin Walters, Rose YuICML 2025
Builds on6
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 82 citations
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 41 citations
- Truth or backpropaganda? An empirical investigation of deep learning theoryMicah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller et al.ICLR 2020 · 36 citations
- Piecewise linear activations substantially shape the loss surfaces of neural networksFengxiang He, Bohan Wang, Dacheng TaoICLR 2020 · 33 citations
- Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU NetworksArsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin GolestaniICLR 2020 · 12 citations
Related papers
- Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature ConnectivityZhanpeng Zhou, Yongyi Yang, Xiaojiang Yang, Junchi Yan et al.NeurIPS 2023 · 56 citations
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 52 citations
- Input Space Mode Connectivity in Deep Neural NetworksJakub Vrábel, Ori Shem-Ur, Yaron Oz, David KruegerICLR 2025
- Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural CollapseArthur Jacot, Peter Súkeník, Zihan Wang, Marco MondelliICLR 2025
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential EquationsJiayao Zhang, Hua Wang, Weijie J. SuNeurIPS 2021 · 9 citations
