LPGD: A General Framework for Backpropagation through Embedded Optimization Layers
Anselm Paulus, Georg Martius, Vít Musil
摘要
Embedding parameterized optimization problems as layers into machine learning architectures serves as a powerful inductive bias. Training such architectures with stochastic gradient descent requires care, as degenerate derivatives of the embedded optimization problem often render the gradients uninformative. We propose Lagrangian Proximal Gradient Descent (LPGD) a flexible framework for training architectures with embedded optimization layers that seamlessly integrates into automatic differentiation libraries. LPGD efficiently computes meaningful replacements of the degenerate optimization layer derivatives by re-running the forward solver oracle on a perturbed input. LPGD captures various previously proposed methods as special cases, while fostering deep links to traditional optimization methods. We theoretically analyze our method and demonstrate on historical and synthetic data that LPGD converges faster than gradient descent even in a differentiable setup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Differentiation Through Black-Box Quadratic Programming SolversConnor W. Magoon, Fengyu Yang, Noam Aigerman, Shahar Z. KovalskyNeurIPS 2025 · 被引用 14 次
- Geometric Algorithms for Neural Combinatorial Optimization with ConstraintsNikolaos Karalias, Akbar Rafiey, Yifei Xu, Zhishang Luo 等NeurIPS 2025 · 被引用 4 次
- SoftJAX & SoftTorch: Empowering Automatic Differentiation Libraries with Informative GradientsAnselm Paulus, Andreas René Geist, Vit Musil, Sebastian Hoffmann 等ICML 2026 · 被引用 3 次
- A Fully First-Order Layer for Differentiable OptimizationZihao Zhao, Kai-Chia Mo, Shing-Hei Ho, Brandon Amos 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper20
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius 等ICLR 2020 · 被引用 341 次
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Implicit Graph Neural NetworksFangda Gu, Heng Chang, Wenwu Zhu, Somayeh Sojoudi 等NeurIPS 2020 · 被引用 188 次
- Smart Predict-and-Optimize for Hard Combinatorial Optimization ProblemsJayanta Mandi, Emir Demirovic, Peter J. Stuckey, Tias GunsAAAI 2020 · 被引用 184 次
- Learning with Differentiable Pertubed OptimizersQuentin Berthet, Mathieu Blondel, Olivier Teboul, Marco Cuturi 等NeurIPS 2020 · 被引用 181 次
相关 Paper
- Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learningAntoine Bambade, Fabian Schramm, Adrien B. Taylor, Justin CarpentierICLR 2024 · 被引用 9 次
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 被引用 34 次
- Alternating Differentiation for Optimization LayersHaixiang Sun, Ye Shi, Jingya Wang, Hoang Duong Tuan 等ICLR 2023 · 被引用 3 次
- Physarum Powered Differentiable Linear Programming Layers and ApplicationsZihang Meng, Sathya N. Ravi, Vikas SinghAAAI 2021 · 被引用 5 次
- ∇-Prox: Differentiable Proximal Algorithm Modeling for Large-Scale OptimizationZeqiang Lai, Kaixuan Wei, Ying Fu, Philipp Härtel 等SIGGRAPH 2023 · 被引用 11 次
