Symmetry Induces Structure and Constraint of Learning
Liu Ziyin
Abstract
Due to common architecture designs, symmetries exist extensively in contemporary neural networks. In this work, we unveil the importance of the loss function symmetries in affecting, if not deciding, the learning behavior of machine learning models. We prove that every mirror-reflection symmetry, with reflection surface , in the loss function leads to the emergence of a constraint on the model parameters : . This constrained solution becomes satisfied when either the weight decay or gradient noise is large. Common instances of mirror symmetries in deep learning include rescaling, rotation, and permutation symmetry. As direct corollaries, we show that rescaling symmetry leads to sparsity, rotation symmetry leads to low rankness, and permutation symmetry leads to homogeneous ensembling. Then, we show that the theoretical framework can explain intriguing phenomena, such as the loss of plasticity and various collapse phenomena in neural networks, and suggest how symmetries can be used to design an elegant algorithm to enforce hard constraints in a differentiable way.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Parameter Symmetry and Noise Equilibrium of Stochastic Gradient DescentLiu Ziyin, Mingze Wang, Hongchao Li, Lei WuNeurIPS 2024 · 23 citations
- Neural Thermodynamics: Entropic Forces in Deep and Universal Representation LearningLiu Ziyin, Yizhou Xu, Isaac L. ChuangNeurIPS 2025 · 11 citations
- Topological obstruction to the training of shallow ReLU neural networksMarco Nurisso, Pierrick Leroy, Francesco VaccarinoNeurIPS 2024 · 6 citations
- Differentiable Sparsity via -Gating: Simple and Versatile Structured PenalizationChris Kolb, Laetitia Frost, Bernd Bischl, David RügamerNeurIPS 2025 · 4 citations
- A universal compression theory for lottery ticket hypothesis and neural scaling lawsHong-Yi Wang, Di Luo, Tomaso Poggio, Isaac L. Chuang et al.ICLR 2026 · 2 citations
Builds on12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
Related papers
- Remove Symmetries to Control Model Expressivity and Improve OptimizationLiu Ziyin, Yizhou Xu, Isaac L. ChuangICLR 2025
- Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of PlasticityAmir Joudaki, Giulia Lanzillotta, Mohammad Samragh, Iman Mirzadeh et al.ICLR 2026 · 5 citations
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro et al.ICML 2021 · 136 citations
- Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural CollapseArthur Jacot, Peter Súkeník, Zihan Wang, Marco MondelliICLR 2025
