Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport
Lingkai Kong, Yuqing Wang, Molei Tao
Abstract
The problem of optimization on Stiefel manifold, i.e., minimizing functions of (not necessarily square) matrices that satisfy orthogonality constraints, has been extensively studied. Yet, a new approach is proposed based on, for the first time, an interplay between thoughtfully designed continuous and discrete dynamics. It leads to a gradient-based optimizer with intrinsically added momentum. This method exactly preserves the manifold structure but does not require additional operation to keep momentum in the changing (co)tangent space, and thus has low computational cost and pleasant accuracy. Its generalization to adaptive learning rates is also demonstrated. Notable performances are observed in practical tasks. For instance, we found that placing orthogonal constraints on attention heads of trained-from-scratch Vision Transformer (Dosovitskiy et al., 2020) could markedly improve its performance, when our optimizer is used, and it is better that each head is made orthogonal within itself but not necessarily to other heads. This optimizer also makes the useful notion of Projection Robust Wasserstein Distance (Paty and Cuturi, 2019; Lin et al., 2020) for high-dim. optimal transport even more effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Fixed Non-negative Orthogonal Classifier: Inducing Zero-mean Neural Collapse with Feature Dimension SeparationHoyong Kim, Kangil KimICLR 2024 · 7 citations
- Quantitative Convergences of Lie Group Momentum OptimizersLingkai Kong, Molei TaoNeurIPS 2024 · 4 citations
- Stiefel Flow Matching for Moment-Constrained Structure ElucidationAustin Henry Cheng, Alston Lo, Kin Long Kelvin Lee, Santiago Miret et al.ICLR 2025 · 2 citations
- Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold MethodAndi Han, Pierre-Louis Poirion, Akiko TakedaICML 2025
Builds on10
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingZizhao Zhang, Han Zhang, Long Zhao, Ting Chen et al.AAAI 2022 · 216 citations
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
Related papers
- Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNsFanchen Bu, Dong Eui ChangAAAI 2022 · 7 citations
- OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit BiasesZixiao Wang, Yifei Shen, Huishuai ZhangICML 2026
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel ManifoldZhizhong Li, Sina Sajadmanesh, Jingtao Li, Lingjuan LyuNeurIPS 2025 · 16 citations
- Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep LearningWu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen et al.ICML 2023 · 13 citations
- Acceleration via silver step-size on Riemannian manifolds with applications to Wasserstein spaceJiyoung Park, Abhishek Roy, Jonathan W. Siegel, Anirban BhattacharyaNeurIPS 2025 · 3 citations
