Orthogonal Over-Parameterized Training
Weiyang Liu, Rongmei Lin, Zhen Liu, James M. Rehg, Liam Paull, Li Xiong, Le Song, Adrian Weller
Abstract
The inductive bias of a neural network is largely determined by the architecture and the training algorithm. To achieve good generalization, how to effectively train a neural network is of great importance. We propose a novel orthogonal over-parameterized training (OPT) framework that can provably minimize the hyperspherical energy which characterizes the diversity of neurons on a hypersphere. By maintaining the minimum hyperspherical energy during training, OPT can greatly improve the empirical generalization. Specifically, OPT fixes the randomly initialized weights of the neurons and learns an orthogonal transformation that applies to these neurons. We consider multiple ways to learn such an orthogonal transformation, including unrolling orthogonalization algorithms, applying orthogonal parameterization, and designing orthogonality-preserving gradient descent. For better scalability, we propose the stochastic OPT which performs orthogonal transformation stochastically for partial dimensions of neurons. Interestingly, OPT reveals that learning a proper coordinate system for neurons is crucial to generalization. We provide some insights on why OPT yields better generalization. Extensive experiments validate the superiority of OPT over the standard training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79d1a610-85f6-4e68-810e-ac36ac658477Cited by top-tier papers25
- Controlling Text-to-Image Diffusion by Orthogonal FinetuningZeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue et al.NeurIPS 2023 · 277 citations
- Parameter-Efficient Orthogonal Finetuning via Butterfly FactorizationWeiyang Liu, Zeju Qiu, Yao Feng, Yuliang Xiu et al.ICLR 2024 · 111 citations
- Towards Principled Disentanglement for Domain GeneralizationHanlin Zhang, Yifan Zhang, Weiyang Liu, Adrian Weller et al.CVPR 2022 · 100 citations
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- SphereFace2: Binary Classification is All You Need for Deep Face RecognitionYandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj et al.ICLR 2022 · 70 citations
Builds on6
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 845 citations
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
- Orthogonal Relation Transforms with Graph Context Modeling for Knowledge Graph EmbeddingYun Tang, Jing Huang, Guangtao Wang, Xiaodong He et al.ACL 2020 · 92 citations
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma et al.ICML 2020 · 57 citations
- Regularizing Neural Networks via Minimizing Hyperspherical EnergyRongmei Lin, Weiyang Liu, Zhen Liu, Chen Feng et al.CVPR 2020
Related papers
- Neural Characteristic Activation Analysis and Geometric Parameterization for ReLU NetworksWenlin Chen, Hong GeNeurIPS 2024 · 5 citations
- Scaling-up Diverse Orthogonal Convolutional Networks by a Paraunitary FrameworkJiahao Su, Wonmin Byeon, Furong HuangICML 2022 · 1 citation
- Reparameterized LLM Training via Orthogonal Equivalence TransformationZeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax et al.NeurIPS 2025 · 11 citations
- Improving Generalization of Deep Neural Networks by Optimum ShiftingYuyan Zhou, Ye Li, Lei Feng, Sheng-Jun HuangAAAI 2025 · 1 citation
- Coordinate Descent on the Orthogonal Group for Recurrent Neural Network TrainingEstelle M. Massart, Vinayak AbrolAAAI 2022 · 13 citations
