Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport
Lingkai Kong, Yuqing Wang, Molei Tao
摘要
The problem of optimization on Stiefel manifold, i.e., minimizing functions of (not necessarily square) matrices that satisfy orthogonality constraints, has been extensively studied. Yet, a new approach is proposed based on, for the first time, an interplay between thoughtfully designed continuous and discrete dynamics. It leads to a gradient-based optimizer with intrinsically added momentum. This method exactly preserves the manifold structure but does not require additional operation to keep momentum in the changing (co)tangent space, and thus has low computational cost and pleasant accuracy. Its generalization to adaptive learning rates is also demonstrated. Notable performances are observed in practical tasks. For instance, we found that placing orthogonal constraints on attention heads of trained-from-scratch Vision Transformer (Dosovitskiy et al., 2020) could markedly improve its performance, when our optimizer is used, and it is better that each head is made orthogonal within itself but not necessarily to other heads. This optimizer also makes the useful notion of Projection Robust Wasserstein Distance (Paty and Cuturi, 2019; Lin et al., 2020) for high-dim. optimal transport even more effective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fixed Non-negative Orthogonal Classifier: Inducing Zero-mean Neural Collapse with Feature Dimension SeparationHoyong Kim, Kangil KimICLR 2024 · 被引用 7 次
- Quantitative Convergences of Lie Group Momentum OptimizersLingkai Kong, Molei TaoNeurIPS 2024 · 被引用 4 次
- Stiefel Flow Matching for Moment-Constrained Structure ElucidationAustin Henry Cheng, Alston Lo, Kin Long Kelvin Lee, Santiago Miret 等ICLR 2025 · 被引用 2 次
- Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold MethodAndi Han, Pierre-Louis Poirion, Akiko TakedaICML 2025
它引用的顶会 Paper10
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingZizhao Zhang, Han Zhang, Long Zhao, Ting Chen 等AAAI 2022 · 被引用 216 次
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 被引用 139 次
相关 Paper
- Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNsFanchen Bu, Dong Eui ChangAAAI 2022 · 被引用 7 次
- OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit BiasesZixiao Wang, Yifei Shen, Huishuai ZhangICML 2026
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel ManifoldZhizhong Li, Sina Sajadmanesh, Jingtao Li, Lingjuan LyuNeurIPS 2025 · 被引用 16 次
- Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep LearningWu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen 等ICML 2023 · 被引用 13 次
- Acceleration via silver step-size on Riemannian manifolds with applications to Wasserstein spaceJiyoung Park, Abhishek Roy, Jonathan W. Siegel, Anirban BhattacharyaNeurIPS 2025 · 被引用 3 次
