Regularizing Neural Networks via Minimizing Hyperspherical Energy
Rongmei Lin, Weiyang Liu, Zhen Liu, Chen Feng, Zhiding Yu, James M. Rehg, Li Xiong, Le Song
摘要
Inspired by the Thomson problem in physics where the distribution of multiple propelling electrons on a unit sphere can be modeled via minimizing some potential energy, hyperspherical energy minimization has demonstrated its potential in regularizing neural networks and improving their generalization power. In this paper, we first study the important role that hyperspherical energy plays in neural network training by analyzing its training dynamics. Then we show that naively minimizing hyperspherical energy suffers from some difficulties due to highly non-linear and non-convex optimization as the space dimensionality becomes higher, therefore limiting the potential to further improve the generalization. To address these problems, we propose the compressive minimum hyperspherical energy (CoMHE) as a more effective regularization for neural networks. Specifically, CoMHE utilizes projection mappings to reduce the dimensionality of neurons and minimizes their hyperspherical energy. According to different designs for the projection mapping, we propose several distinct yet well-performing variants and provide some theoretical guarantees to justify their effectiveness. Our experiments show that CoMHE consistently outperforms existing regularization methods, and can be easily applied to different neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Controlling Text-to-Image Diffusion by Orthogonal FinetuningZeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue 等NeurIPS 2023 · 被引用 277 次
- SphereFace2: Binary Classification is All You Need for Deep Face RecognitionYandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj 等ICLR 2022 · 被引用 70 次
- Angular Visual HardnessBeidi Chen, Weiyang Liu, Zhiding Yu, Jan Kautz 等ICML 2020 · 被引用 57 次
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah 等CVPR 2022 · 被引用 36 次
- Maximum Class Separation as Inductive Bias in One MatrixTejaswi Kasarla, Gertjan J. Burghouts, Max van Spengler, Elise van der Pol 等NeurIPS 2022 · 被引用 29 次
它引用的顶会 Paper1
相关 Paper
- MMA Regularization: Decorrelating Weights of Neural Networks by Maximizing the Minimal AnglesZhennan Wang, Canqun Xiang, Wenbin Zou, Chen XuNeurIPS 2020 · 被引用 25 次
- Generalizing and Decoupling Neural Collapse via Hyperspherical Uniformity GapWeiyang Liu, Longhui Yu, Adrian Weller, Bernhard SchölkopfICLR 2023 · 被引用 3 次
- Connecting Sphere Manifolds Hierarchically for RegularizationDamien Scieur, Youngsung KimICML 2021 · 被引用 3 次
- On Learning Deep O(n)-Equivariant HyperspheresPavlo Melnyk, Michael Felsberg, Mårten Wadenbäck, Andreas Robinson 等ICML 2024
- nGPT: Normalized Transformer with Representation Learning on the HypersphereIlya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris GinsburgICLR 2025
