Do Deep Neural Network Solutions Form a Star Domain?
Ankit Sonthalia, Alexander Rubinstein, Ehsan Abbasnejad, Seong Joon Oh
Abstract
It has recently been conjectured that neural network solution sets reachable via stochastic gradient descent (SGD) are convex, considering permutation invariances (Entezari et al., 2021). This means that a linear path can connect two independent solutions with low loss, given the weights of one of the models are appropriately permuted. However, current methods to test this theory often require very wide networks to succeed (Ainsworth et al., 2022;Benzing et al., 2022). In this work, we conjecture that more generally, the SGD solution set is a star domain that contains a star model that is linearly connected to all the other solutions via paths with low loss values, modulo permutations. We propose the Starlight algorithm that finds a star model of a given learning task. We validate our claim by showing that this star model is linearly connected with other independently found solutions. As an additional benefit of our study, we demonstrate better uncertainty estimates on Bayesian Model Averaging over the obtained star domain. Further, we demonstrate star models as potential substitutes for model ensembles. Our code is available at https://github.com/aktsonthalia/starlight.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e3fbe65-ebf8-4496-a394-a1642d87a2eeCited by top-tier papers5
- On Linear Mode Connectivity of Mixture-of-Experts ArchitecturesViet-Hoang Tran, Van-Hoan Trinh, Khanh Vinh Bui, Tan M. NguyenNeurIPS 2025 · 9 citations
- Flat Channels to Infinity in Neural Loss LandscapesFlavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner et al.NeurIPS 2025 · 6 citations
- On The Surprising Effectiveness of a Single Global Merging in Decentralized LearningTongtian Zhu, Tianyu Zhang, Mingze Wang, Zhanpeng Zhou et al.ICLR 2026 · 2 citations
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial ConditionsGül Sena Altintas, Devin Kwok, Colin Raffel, David RolnickICML 2025
- Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode ConnectivityViet Hoang Tran, VINH KHANH BUI, Van-Hoan Trinh, Ngoc Tan Lai et al.ICML 2026
Builds on14
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- Bridging Mode Connectivity in Loss Landscapes and Adversarial RobustnessPu Zhao, Pin-Yu Chen, Payel Das, Karthikeyan Natesan Ramamurthy et al.ICLR 2020 · 213 citations
Related papers
- Linear Mode Connectivity between Multiple Models modulo Permutation SymmetriesAkira Ito, Masanori Yamada, Atsutoshi KumagaiICML 2025
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 32 citations
- Learning Neural Network SubspacesMitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi et al.ICML 2021 · 101 citations
- Characterizing the Loss Landscape in Non-Negative Matrix FactorizationJohan Bjorck, Anmol Kabra, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 4 citations
- Loss Surface Simplexes for Mode Connecting Volumes and Fast EnsemblingGregory W. Benton, Wesley J. Maddox, Sanae Lotfi, Andrew Gordon WilsonICML 2021 · 88 citations
