Efficient Riemannian Meta-Optimization by Implicit Differentiation
Xiaomeng Fan, Yuwei Wu, Zhi Gao, Yunde Jia, Mehrtash Harandi
Abstract
To solve optimization problems with nonlinear constrains, the recently developed Riemannian meta-optimization methods show promise, which train neural networks as an optimizer to perform optimization on Riemannian manifolds. A key challenge is the heavy computational and memory burdens, because computing the meta-gradient with respect to the optimizer involves a series of time-consuming derivatives, and stores large computation graphs in memory. In this paper, we propose an efficient Riemannian meta-optimization method that decouples the complex computation scheme from the meta-gradient. We derive Riemannian implicit differentiation to compute the meta-gradient by establishing a link between Riemannian optimization and the implicit function theorem. As a result, the updating our optimizer is only related to the final two iterations, which in turn speeds up our method and reduces the memory footprint significantly. We theoretically study the computational load and memory footprint of our method for long optimization trajectories, and conduct an empirical study to demonstrate the benefits of the proposed method. Evaluations of three optimization problems on different Riemannian manifolds show that our method achieves state-of-the-art performance in terms of the convergence speed and the quality of optima.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Training Stronger Baselines for Learning to OptimizeTianlong Chen, Weiyi Zhang, Jingyang Zhou, Shiyu Chang et al.NeurIPS 2020 · 61 citations
- Learning a Gradient-free Riemannian Optimizer on Tangent SpacesXiaomeng Fan, Zhi Gao, Yuwei Wu, Yunde Jia et al.AAAI 2021 · 8 citations
- Learning to Optimize on SPD ManifoldsZhi Gao, Yuwei Wu, Yunde Jia, Mehrtash HarandiCVPR 2020
- AutoDO: Robust AutoAugment for Biased Data With Label Noise via Scalable Probabilistic Implicit DifferentiationDenis A. Gudovskiy, Luca Rigazio, Shun Ishizaka, Kazuki Kozuka et al.CVPR 2021
Related papers
- Decentralized Riemannian Conjugate Gradient Method on the Stiefel ManifoldJun Chen, Haishan Ye, Mengmeng Wang, Tianxin Huang et al.ICLR 2024 · 21 citations
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 10 citations
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
- Feedback Gradient Descent: Efficient and Stable Optimization with Orthogonality for DNNsFanchen Bu, Dong Eui ChangAAAI 2022 · 7 citations
- On the Iteration Complexity of Hypergradient ComputationRiccardo Grazzi, Luca Franceschi, Massimiliano Pontil, Saverio SalzoICML 2020 · 241 citations
