An Embarrassingly Simple Way to Optimize Orthogonal Matrices at Scale
Adrián Javaloy, Antonio Vergari
摘要
Orthogonality constraints are ubiquitous in robust and probabilistic machine learning. Unfortunately, current optimizers are computationally expensive and do not scale to problems with hundreds or thousands of constraints. One notable exception is the Landing algorithm (Ablin et al., 2024) which, however comes at the expense of temporarily relaxing orthogonality. In this work, we revisit and improve on the ideas behind Landing, enabling the inclusion of modern adaptive optimizers while ensuring that orthogonal constraints are effectively met. Remarkably, these improvements come at little to no cost, and reduce the number of required hyperparemeters. Our algorithm POGO is fast and GPU-friendly, consisting of only 5 matrix products , and in practice maintains orthogonality at all times. On several challenging benchmarks, POGO greatly outperforms recent optimizers and shows it can optimize problems with thousands of orthogonal matrices in minutes while alternatives would take hours. As such, POGO sets a milestone to finally exploit orthogonality constraints in ML at scale. A public PyTorch implementation of POGO is available at https://github.com/adrianjav/pogo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- RotoGrad: Gradient Homogenization in Multitask LearningAdrián Javaloy, Isabel ValeraICLR 2022 · 被引用 114 次
- A Compositional Atlas of Tractable Circuit Operations for Probabilistic InferenceAntonio Vergari, YooJung Choi, Anji Liu, Stefano Teso 等NeurIPS 2021 · 被引用 112 次
- projUNN: efficient method for training deep networks with unitary matricesBobak Toussi Kiani, Randall Balestriero, Yann LeCun, Seth LloydNeurIPS 2022 · 被引用 42 次
- Subtractive Mixture Models via Squaring: Representation and LearningLorenzo Loconte, Aleksanteri M. Sladek, Stefan Mengel, Martin Trapp 等ICLR 2024 · 被引用 42 次
- VectorAdam for Rotation Equivariant Geometry OptimizationSelena Ling, Nicholas Sharp, Alec JacobsonNeurIPS 2022 · 被引用 28 次
相关 Paper
- Distributed Retraction-Free and Communication-Efficient Optimization on the Stiefel ManifoldYilong Song, Peijin Li, Bin Gao, Kun YuanICML 2025
- An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN ArchitecturesThibaut Boissin, Franck Mamalet, Thomas Fel, Agustin Martin Picard 等ICML 2025
- Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold MethodAndi Han, Pierre-Louis Poirion, Akiko TakedaICML 2025
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 被引用 17 次
- Hom-PGD: Fast Reparameterized Optimization over Non-convex Ball-Homeomorphic SetChenghao Liu, Enming Liang, Minghua ChenICML 2026
