Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNs
Fanchen Bu, Dong Eui Chang
摘要
The optimization with orthogonality has been shown useful in training deep neural networks (DNNs). To impose orthogonality on DNNs, both computational efficiency and stability are important. However, existing methods utilizing Riemannian optimization or hard constraints can only ensure stability while those using soft constraints can only improve efficiency. In this paper, we propose a novel method, named Feedback Gradient Descent (FGD), to our knowledge, the first work showing high efficiency and stability simultaneously. FGD induces orthogonality based on the simple yet indispensable Euler discretization of a continuous-time dynamical system on the tangent bundle of the Stiefel manifold. In particular, inspired by a numerical integration method on manifolds called Feedback Integrators, we propose to instantiate it on the tangent bundle of the Stiefel manifold for the first time. In our extensive image classification experiments, FGD comprehensively outperforms the existing state-of-the-art methods in terms of accuracy, efficiency, and stability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Toward Equation of Motion for Deep Neural Networks: Continuous-time Gradient Descent and Discretization Error AnalysisTaiki MiyagawaNeurIPS 2022 · 被引用 14 次
- Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal TransportLingkai Kong, Yuqing Wang, Molei TaoICLR 2023
它引用的顶会 Paper7
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 被引用 139 次
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 被引用 137 次
- Accelerating SGD with momentum for over-parameterized learningChaoyue Liu, Mikhail BelkinICLR 2020 · 被引用 93 次
- Convolutional Normalization: Improving Deep Convolutional Network Robustness and TrainingSheng Liu, Xiao Li, Yuexiang Zhai, Chong You 等NeurIPS 2021 · 被引用 30 次
- Eigenvalue Normalized Recurrent Neural Networks for Short Term MemoryKyle Helfrich, Qiang YeAAAI 2020 · 被引用 8 次
相关 Paper
- Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold MethodAndi Han, Pierre-Louis Poirion, Akiko TakedaICML 2025
- Decentralized Riemannian Algorithm for Nonconvex Minimax ProblemsXidong Wu, Zhengmian Hu, Heng HuangAAAI 2023 · 被引用 15 次
- Controllable Orthogonalization in Training DNNsLei Huang, Li Liu, Fan Zhu, Diwen Wan 等CVPR 2020
- Efficient Riemannian Meta-Optimization by Implicit DifferentiationXiaomeng Fan, Yuwei Wu, Zhi Gao, Yunde Jia 等AAAI 2022 · 被引用 3 次
- Distributed Retraction-Free and Communication-Efficient Optimization on the Stiefel ManifoldYilong Song, Peijin Li, Bin Gao, Kun YuanICML 2025
