Transformer-Based Learned Optimization
Erik Gärtner, Luke Metz, Mykhaylo Andriluka, C. Daniel Freeman, Cristian Sminchisescu
摘要
We propose a new approach to learned optimization where we represent the computation of an optimizer's update step using a neural network. The parameters of the optimizer are then learned by training on a set of optimization tasks with the objective to perform minimization efficiently. Our innovation is a new neural network architecture, Optimus, for the learned optimizer inspired by the classic BFGS algorithm. As in BFGS, we estimate a preconditioning matrix as a sum of rank-one updates but use a Transformerbased neural network to predict these updates jointly with the step length and direction. In contrast to several recent learned optimization-based approaches [24, 27] , our formulation allows for conditioning across the dimensions of the parameter space of the target problem while remaining applicable to optimization tasks of variable dimensionality without retraining. We demonstrate the advantages of our approach on a benchmark composed of objective functions traditionally used for the evaluation of optimization algorithms, as well as on the real world-task of physics-based visual reconstruction of articulated 3d human motion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- B2Opt: Learning to Optimize Black-box Optimization with Little BudgetXiaobin Li, Kai Wu, Xiaoyu Zhang, Handing WangAAAI 2025 · 被引用 23 次
- Can Learned Optimization Make Reinforcement Learning Less Difficult?Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson 等NeurIPS 2024 · 被引用 18 次
- Mnemosyne: Learning to Train Transformers with TransformersDeepali Jain, Krzysztof Marcin Choromanski, Kumar Avinava Dubey, Sumeet Singh 等NeurIPS 2023 · 被引用 15 次
- Differentiable Approximations for Distance QueriesAhmed Abdelkader, David M. MountSODA 2025
- L-SR1: Learned Symmetric-Rank-One PreconditioningGal Lifshitz, Shahar Zuler, Ori Fouks, Dan RavivICML 2026
它引用的顶会 Paper9
- Neural monocular 3D human motion capture with physical awarenessSoshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick Pérez 等SIGGRAPH 2021 · 被引用 107 次
- Physics-based Human Motion Estimation and Synthesis from VideosKevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo 等ICCV 2021 · 被引用 102 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- Differentiable Dynamics for Articulated 3d Human Motion ReconstructionErik Gärtner, Mykhaylo Andriluka, Erwin Coumans, Cristian SminchisescuCVPR 2022 · 被引用 33 次
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun 等NeurIPS 2021 · 被引用 27 次
相关 Paper
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon 等ICLR 2026 · 被引用 10 次
- Celo2: Towards Learned Optimization Free LunchAbhinav Moudgil, Boris Knyazev, Eugene BelilovskyICLR 2026 · 被引用 1 次
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin 等ICML 2023 · 被引用 18 次
- Learning to Learn by Zeroth-Order OracleYangjun Ruan, Yuanhao Xiong, Sashank J. Reddi, Sanjiv Kumar 等ICLR 2020 · 被引用 21 次
- Potential Based Diffusion Motion PlanningYunhao Luo, Chen Sun, Joshua B. Tenenbaum, Yilun DuICML 2024 · 被引用 43 次
