LPGD: A General Framework for Backpropagation through Embedded Optimization Layers
Anselm Paulus, Georg Martius, Vít Musil
Abstract
Embedding parameterized optimization problems as layers into machine learning architectures serves as a powerful inductive bias. Training such architectures with stochastic gradient descent requires care, as degenerate derivatives of the embedded optimization problem often render the gradients uninformative. We propose Lagrangian Proximal Gradient Descent (LPGD) a flexible framework for training architectures with embedded optimization layers that seamlessly integrates into automatic differentiation libraries. LPGD efficiently computes meaningful replacements of the degenerate optimization layer derivatives by re-running the forward solver oracle on a perturbed input. LPGD captures various previously proposed methods as special cases, while fostering deep links to traditional optimization methods. We theoretically analyze our method and demonstrate on historical and synthetic data that LPGD converges faster than gradient descent even in a differentiable setup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Differentiation Through Black-Box Quadratic Programming SolversConnor W. Magoon, Fengyu Yang, Noam Aigerman, Shahar Z. KovalskyNeurIPS 2025 · 14 citations
- Geometric Algorithms for Neural Combinatorial Optimization with ConstraintsNikolaos Karalias, Akbar Rafiey, Yifei Xu, Zhishang Luo et al.NeurIPS 2025 · 4 citations
- SoftJAX & SoftTorch: Empowering Automatic Differentiation Libraries with Informative GradientsAnselm Paulus, Andreas René Geist, Vit Musil, Sebastian Hoffmann et al.ICML 2026 · 3 citations
- A Fully First-Order Layer for Differentiable OptimizationZihao Zhao, Kai-Chia Mo, Shing-Hei Ho, Brandon Amos et al.ICML 2026 · 1 citation
Builds on20
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius et al.ICLR 2020 · 341 citations
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Implicit Graph Neural NetworksFangda Gu, Heng Chang, Wenwu Zhu, Somayeh Sojoudi et al.NeurIPS 2020 · 188 citations
- Smart Predict-and-Optimize for Hard Combinatorial Optimization ProblemsJayanta Mandi, Emir Demirovic, Peter J. Stuckey, Tias GunsAAAI 2020 · 184 citations
- Learning with Differentiable Pertubed OptimizersQuentin Berthet, Mathieu Blondel, Olivier Teboul, Marco Cuturi et al.NeurIPS 2020 · 181 citations
Related papers
- Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learningAntoine Bambade, Fabian Schramm, Adrien B. Taylor, Justin CarpentierICLR 2024 · 9 citations
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 34 citations
- Alternating Differentiation for Optimization LayersHaixiang Sun, Ye Shi, Jingya Wang, Hoang Duong Tuan et al.ICLR 2023 · 3 citations
- Physarum Powered Differentiable Linear Programming Layers and ApplicationsZihang Meng, Sathya N. Ravi, Vikas SinghAAAI 2021 · 5 citations
- ∇-Prox: Differentiable Proximal Algorithm Modeling for Large-Scale OptimizationZeqiang Lai, Kaixuan Wei, Ying Fu, Philipp Härtel et al.SIGGRAPH 2023 · 11 citations
