Coordinate Descent on the Orthogonal Group for Recurrent Neural Network Training
Estelle M. Massart, Vinayak Abrol
Abstract
We address the poor scalability of learning algorithms for orthogonal recurrent neural networks via the use of stochastic coordinate descent on the orthogonal group, leading to a cost per iteration that increases linearly with the number of recurrent states. This contrasts with the cubic dependency of typical feasible algorithms such as stochastic Riemannian gradient descent, which prohibits the use of big network architectures. Coordinate descent rotates successively two columns of the recurrent matrix. When the coordinate (i.e., indices of rotated columns) is selected uniformly at random at each iteration, we prove convergence of the algorithm under standard assumptions on the loss function, stepsize and minibatch noise. In addition, we numerically show that the Riemannian gradient has an approximately sparse structure. Leveraging this observation, we propose a variant of our proposed algorithm that relies on the Gauss-Southwell coordinate selection rule. Experiments on a benchmark recurrent neural network training problem show that the proposed approach is a very promising step towards the training of orthogonal recurrent neural networks with big architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b580938-37d5-4fe1-a967-d745eed0c9bcCited by top-tier papers4
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 10 citations
- Times2D: Multi-Period Decomposition and Derivative Mapping for General Time Series ForecastingReza Nematirad, Anil Pahwa, Balasubramaniam NatarajanAAAI 2025 · 9 citations
- A Block Coordinate Descent Method for Nonsmooth Composite Optimization under Orthogonality ConstraintsGanzhao YuanICLR 2026 · 5 citations
- OT4P: Unlocking Effective Orthogonal Group Path for Permutation RelaxationYaming Guo, Chen Zhu, Hengshu Zhu, Tieru WuNeurIPS 2024 · 1 citation
Builds on2
Related papers
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
- Stochastic Flows and Geometric Optimization on the Orthogonal GroupKrzysztof Choromanski, David Cheikhi, Jared Davis, Valerii Likhosherstov et al.ICML 2020 · 7 citations
- Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold MethodAndi Han, Pierre-Louis Poirion, Akiko TakedaICML 2025
- Orthogonal Over-Parameterized TrainingWeiyang Liu, Rongmei Lin, Zhen Liu, James M. Rehg et al.CVPR 2021
- projUNN: efficient method for training deep networks with unitary matricesBobak Toussi Kiani, Randall Balestriero, Yann LeCun, Seth LloydNeurIPS 2022 · 42 citations
