Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNs
Fanchen Bu, Dong Eui Chang
Abstract
The optimization with orthogonality has been shown useful in training deep neural networks (DNNs). To impose orthogonality on DNNs, both computational efficiency and stability are important. However, existing methods utilizing Riemannian optimization or hard constraints can only ensure stability while those using soft constraints can only improve efficiency. In this paper, we propose a novel method, named Feedback Gradient Descent (FGD), to our knowledge, the first work showing high efficiency and stability simultaneously. FGD induces orthogonality based on the simple yet indispensable Euler discretization of a continuous-time dynamical system on the tangent bundle of the Stiefel manifold. In particular, inspired by a numerical integration method on manifolds called Feedback Integrators, we propose to instantiate it on the tangent bundle of the Stiefel manifold for the first time. In our extensive image classification experiments, FGD comprehensively outperforms the existing state-of-the-art methods in terms of accuracy, efficiency, and stability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Toward Equation of Motion for Deep Neural Networks: Continuous-time Gradient Descent and Discretization Error AnalysisTaiki MiyagawaNeurIPS 2022 · 14 citations
- Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal TransportLingkai Kong, Yuqing Wang, Molei TaoICLR 2023
Builds on7
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 137 citations
- Accelerating SGD with momentum for over-parameterized learningChaoyue Liu, Mikhail BelkinICLR 2020 · 93 citations
- Convolutional Normalization: Improving Deep Convolutional Network Robustness and TrainingSheng Liu, Xiao Li, Yuexiang Zhai, Chong You et al.NeurIPS 2021 · 30 citations
- Eigenvalue Normalized Recurrent Neural Networks for Short Term MemoryKyle Helfrich, Qiang YeAAAI 2020 · 8 citations
Related papers
- Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold MethodAndi Han, Pierre-Louis Poirion, Akiko TakedaICML 2025
- Decentralized Riemannian Algorithm for Nonconvex Minimax ProblemsXidong Wu, Zhengmian Hu, Heng HuangAAAI 2023 · 15 citations
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan et al.CVPR 2020
- Efficient Riemannian Meta-Optimization by Implicit DifferentiationXiaomeng Fan, Yuwei Wu, Zhi Gao, Yunde Jia et al.AAAI 2022 · 3 citations
- Distributed Retraction-Free and Communication-Efficient Optimization on the Stiefel ManifoldYilong Song, Peijin Li, Bin Gao, Kun YuanICML 2025
