Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learning
Antoine Bambade, Fabian Schramm, Adrien B. Taylor, Justin Carpentier
Abstract
Optimization layers within neural network architectures have become increasingly popular for their ability to solve a wide range of machine learning tasks and to model domain-specific knowledge. However, designing optimization layers requires careful consideration as the underlying optimization problems might be infeasible during training. Motivated by applications in learning, control and robotics, this work focuses on convex quadratic programming (QP) layers. The specific structure of this type of optimization layer can be efficiently exploited for faster computations while still allowing rich modeling capabilities. We leverage primal-dual augmented Lagrangian techniques for computing derivatives of both feasible and infeasible QP solutions. More precisely, we propose a unified approach that tackles the differentiability of the closest feasible QP solutions in a classical ℓ 2 sense. We then harness this approach to enrich the expressive capabilities of existing QP layers. More precisely, we show how differentiating through infeasible QPs during training enables to drive towards feasibility at test time a new range of QP layers. These layers notably demonstrate superior predictive performance in some conventional learning tasks. Additionally, we present alternative formulations that enhance numerical robustness, speed, and accuracy for training such layers. Along with these contributions, we provide an open-source C++ software package called QPLayer for differentiating feasible and infeasible convex QPs and which can be interfaced with modern learning frameworks.
Published as a conference paper at ICLR 2024 • In Section 3.4 we provide efficient ways to compute the Jacobian ∂x ⋆ (θ) ∂θ in forward and backward automatic differentiation modes.
• In Section 4 we demonstrate how the approach enables dealing with possibly infeasible QP(θ) during training, while converging for test time to a feasible layer. We illustrate how it allows to train a broader range of QP layers (e.g., learning QPs that are not generically feasible). More precisely, we will show how to drive towards feasibility at test time the QP layer provided in Figure 3. Learning A t (in red) is not obvious since nothing guarantees a priori that the fixed equality constraint vector (of ones) lies in the range space of A t . We will see that learning such layer notably provides better predictive power for some classic learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c4adbd1-8363-48d5-9fa0-0fc0b7d88c5eCited by top-tier papers6
- Differentiation Through Black-Box Quadratic Programming SolversConnor W. Magoon, Fengyu Yang, Noam Aigerman, Shahar Z. KovalskyNeurIPS 2025 · 14 citations
- Differentiable Model Predictive Control on the GPUEmre Adabag, Marcus Greiff, John Subosits, Thomas Jonathan LewICLR 2026 · 13 citations
- A Fully First-Order Layer for Differentiable OptimizationZihao Zhao, Kai-Chia Mo, Shing-Hei Ho, Brandon Amos et al.ICML 2026 · 1 citation
- A Penalty Approach For Differentiation Through Black-box Quadratic Programming SolversYuxuan Linghu, Zhiyuan Liu, Qi DengICML 2026 · 1 citation
- HONet: Data-Efficient Learning for Exact Cover Tasks via Hypergraph OptimizationPengyang Huang, Zirui Zhuang, Haifeng Sun, Qi Qi et al.ICML 2026
Builds on11
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Implicit Surface Representations As Layers in Neural NetworksMateusz Michalkiewicz, Jhony Kaesemodel Pontes, Dominic Jack, Mahsa Baktashmotlagh et al.ICCV 2019 · 298 citations
- Is Attention Better Than Matrix Decomposition?Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li et al.ICLR 2021 · 171 citations
- JFB: Jacobian-Free Backpropagation for Implicit NetworksSamy Wu Fung, Howard Heaton, Qiuwei Li, Daniel McKenzie et al.AAAI 2022 · 123 citations
- On Training Implicit ModelsZhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang et al.NeurIPS 2021 · 111 citations
Related papers
- BPQP: A Differentiable Convex Optimization Framework for Efficient End-to-End LearningJianming Pan, Zeqi Ye, Xiao Yang, Xu Yang et al.NeurIPS 2024 · 18 citations
- LPGD: A General Framework for Backpropagation through Embedded Optimization LayersAnselm Paulus, Georg Martius, Vít MusilICML 2024 · 5 citations
- Pinet: Optimizing hard-constrained neural networks with orthogonal projection layersPanagiotis D. Grontas, Antonio Terpin, Efe C. Balta, Raffaello D'Andrea et al.ICLR 2026 · 22 citations
- QPKO: Differentiable QP-Embedded Deep Koopman Framework for Modeling Nonlinear SystemsRunze Tian, Peng KouICML 2026
- Alternating Differentiation for Optimization LayersHaixiang Sun, Ye Shi, Jingya Wang, Hoang Duong Tuan et al.ICLR 2023 · 3 citations
