Regularizing Neural Networks via Minimizing Hyperspherical Energy
Rongmei Lin, Weiyang Liu, Zhen Liu, Chen Feng, Zhiding Yu, James M. Rehg, Li Xiong, Le Song
Abstract
Inspired by the Thomson problem in physics where the distribution of multiple propelling electrons on a unit sphere can be modeled via minimizing some potential energy, hyperspherical energy minimization has demonstrated its potential in regularizing neural networks and improving their generalization power. In this paper, we first study the important role that hyperspherical energy plays in neural network training by analyzing its training dynamics. Then we show that naively minimizing hyperspherical energy suffers from some difficulties due to highly non-linear and non-convex optimization as the space dimensionality becomes higher, therefore limiting the potential to further improve the generalization. To address these problems, we propose the compressive minimum hyperspherical energy (CoMHE) as a more effective regularization for neural networks. Specifically, CoMHE utilizes projection mappings to reduce the dimensionality of neurons and minimizes their hyperspherical energy. According to different designs for the projection mapping, we propose several distinct yet well-performing variants and provide some theoretical guarantees to justify their effectiveness. Our experiments show that CoMHE consistently outperforms existing regularization methods, and can be easily applied to different neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abb75271-10c5-4df2-a3e7-99697465ffbdCited by top-tier papers18
- Controlling Text-to-Image Diffusion by Orthogonal FinetuningZeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue et al.NeurIPS 2023 · 277 citations
- SphereFace2: Binary Classification is All You Need for Deep Face RecognitionYandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj et al.ICLR 2022 · 70 citations
- Angular Visual HardnessBeidi Chen, Weiyang Liu, Zhiding Yu, Jan Kautz et al.ICML 2020 · 57 citations
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah et al.CVPR 2022 · 36 citations
- Maximum Class Separation as Inductive Bias in One MatrixTejaswi Kasarla, Gertjan J. Burghouts, Max van Spengler, Elise van der Pol et al.NeurIPS 2022 · 29 citations
Builds on1
Related papers
- MMA Regularization: Decorrelating Weights of Neural Networks by Maximizing the Minimal AnglesZhennan Wang, Canqun Xiang, Wenbin Zou, Chen XuNeurIPS 2020 · 25 citations
- Generalizing and Decoupling Neural Collapse via Hyperspherical Uniformity GapWeiyang Liu, Longhui Yu, Adrian Weller, Bernhard SchölkopfICLR 2023 · 3 citations
- Connecting Sphere Manifolds Hierarchically for RegularizationDamien Scieur, Youngsung KimICML 2021 · 3 citations
- On Learning Deep O(n)-Equivariant HyperspheresPavlo Melnyk, Michael Felsberg, Mårten Wadenbäck, Andreas Robinson et al.ICML 2024
- nGPT: Normalized Transformer with Representation Learning on the HypersphereIlya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris GinsburgICLR 2025
