Orthogonal Over-Parameterized Training
Weiyang Liu, Rongmei Lin, Zhen Liu, James M. Rehg, Liam Paull, Li Xiong, Le Song, Adrian Weller
摘要
The inductive bias of a neural network is largely determined by the architecture and the training algorithm. To achieve good generalization, how to effectively train a neural network is of great importance. We propose a novel orthogonal over-parameterized training (OPT) framework that can provably minimize the hyperspherical energy which characterizes the diversity of neurons on a hypersphere. By maintaining the minimum hyperspherical energy during training, OPT can greatly improve the empirical generalization. Specifically, OPT fixes the randomly initialized weights of the neurons and learns an orthogonal transformation that applies to these neurons. We consider multiple ways to learn such an orthogonal transformation, including unrolling orthogonalization algorithms, applying orthogonal parameterization, and designing orthogonality-preserving gradient descent. For better scalability, we propose the stochastic OPT which performs orthogonal transformation stochastically for partial dimensions of neurons. Interestingly, OPT reveals that learning a proper coordinate system for neurons is crucial to generalization. We provide some insights on why OPT yields better generalization. Extensive experiments validate the superiority of OPT over the standard training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Controlling Text-to-Image Diffusion by Orthogonal FinetuningZeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue 等NeurIPS 2023 · 被引用 277 次
- Parameter-Efficient Orthogonal Finetuning via Butterfly FactorizationWeiyang Liu, Zeju Qiu, Yao Feng, Yuliang Xiu 等ICLR 2024 · 被引用 111 次
- Towards Principled Disentanglement for Domain GeneralizationHanlin Zhang, Yifan Zhang, Weiyang Liu, Adrian Weller 等CVPR 2022 · 被引用 100 次
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li 等NeurIPS 2020 · 被引用 99 次
- SphereFace2: Binary Classification is All You Need for Deep Face RecognitionYandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj 等ICLR 2022 · 被引用 70 次
它引用的顶会 Paper6
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 被引用 845 次
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 被引用 139 次
- Orthogonal Relation Transforms with Graph Context Modeling for Knowledge Graph EmbeddingYun Tang, Jing Huang, Guangtao Wang, Xiaodong He 等ACL 2020 · 被引用 92 次
- Deep Isometric Learning for Visual RecognitionHaozhi Qi, Chong You, Xiaolong Wang, Yi Ma 等ICML 2020 · 被引用 57 次
- Regularizing Neural Networks via Minimizing Hyperspherical EnergyRongmei Lin, Weiyang Liu, Zhen Liu, Chen Feng 等CVPR 2020
相关 Paper
- Neural Characteristic Activation Analysis and Geometric Parameterization for ReLU NetworksWenlin Chen, Hong GeNeurIPS 2024 · 被引用 5 次
- Scaling-up Diverse Orthogonal Convolutional Networks by a Paraunitary FrameworkJiahao Su, Wonmin Byeon, Furong HuangICML 2022 · 被引用 1 次
- Reparameterized LLM Training via Orthogonal Equivalence TransformationZeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax 等NeurIPS 2025 · 被引用 11 次
- Improving Generalization of Deep Neural Networks by Optimum ShiftingYuyan Zhou, Ye Li, Lei Feng, Sheng-Jun HuangAAAI 2025 · 被引用 1 次
- Coordinate Descent on the Orthogonal Group for Recurrent Neural Network TrainingEstelle M. Massart, Vinayak AbrolAAAI 2022 · 被引用 13 次
