Lune

ICLR2024Top-tier venue

Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learning

Antoine Bambade, Fabian Schramm, Adrien B. Taylor, Justin Carpentier

2024Year
9Citations
6Top-tier citations

Abstract

Optimization layers within neural network architectures have become increasingly popular for their ability to solve a wide range of machine learning tasks and to model domain-specific knowledge. However, designing optimization layers requires careful consideration as the underlying optimization problems might be infeasible during training. Motivated by applications in learning, control and robotics, this work focuses on convex quadratic programming (QP) layers. The specific structure of this type of optimization layer can be efficiently exploited for faster computations while still allowing rich modeling capabilities. We leverage primal-dual augmented Lagrangian techniques for computing derivatives of both feasible and infeasible QP solutions. More precisely, we propose a unified approach that tackles the differentiability of the closest feasible QP solutions in a classical ℓ 2 sense. We then harness this approach to enrich the expressive capabilities of existing QP layers. More precisely, we show how differentiating through infeasible QPs during training enables to drive towards feasibility at test time a new range of QP layers. These layers notably demonstrate superior predictive performance in some conventional learning tasks. Additionally, we present alternative formulations that enhance numerical robustness, speed, and accuracy for training such layers. Along with these contributions, we provide an open-source C++ software package called QPLayer for differentiating feasible and infeasible convex QPs and which can be interfaced with modern learning frameworks.

Published as a conference paper at ICLR 2024 • In Section 3.4 we provide efficient ways to compute the Jacobian ∂x ⋆ (θ) ∂θ in forward and backward automatic differentiation modes.

• In Section 4 we demonstrate how the approach enables dealing with possibly infeasible QP(θ) during training, while converging for test time to a feasible layer. We illustrate how it allows to train a broader range of QP layers (e.g., learning QPs that are not generically feasible). More precisely, we will show how to drive towards feasibility at test time the QP layer provided in Figure 3. Learning A t (in red) is not obvious since nothing guarantees a priori that the fixed equality constraint vector (of ones) lies in the range space of A t . We will see that learning such layer notably provides better predictive power for some classic learning tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6c4adbd1-8363-48d5-9fa0-0fc0b7d88c5e

Cited by top-tier papers6

Ask how each one uses it

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines