Learning Energy Networks with Generalized Fenchel-Young Losses
Mathieu Blondel, Felipe Llinares-López, Robert Dadashi, Léonard Hussenot, Matthieu Geist
Abstract
Energy-based models, a.k.a. energy networks, perform inference by optimizing an energy function, typically parametrized by a neural network. This allows one to capture potentially complex relationships between inputs and outputs. To learn the parameters of the energy function, the solution to that optimization problem is typically fed into a loss function. The key challenge for training energy networks lies in computing loss gradients, as this typically requires argmin/argmax differentiation. In this paper, building upon a generalized notion of conjugate function, which replaces the usual bilinear pairing with a general energy function, we propose generalized Fenchel-Young losses, a natural loss construction for learning energy networks. Our losses enjoy many desirable properties and their gradients can be computed efficiently without argmin/argmax differentiation. We also prove the calibration of their excess risk in the case of linear-concave energies. We demonstrate our losses on multilabel classification and imitation learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Learning with Fitzpatrick LossesSeta Rakotomandimby, Jean-Philippe Chancelier, Michel De Lara, Mathieu BlondelNeurIPS 2024 · 6 citations
- LPGD: A General Framework for Backpropagation through Embedded Optimization LayersAnselm Paulus, Georg Martius, Vít MusilICML 2024 · 5 citations
- Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young LossesYuzhou Cao, Han Bao, Lei Feng, Bo AnNeurIPS 2025 · 4 citations
- Non-Stationary Online Structured Prediction with Surrogate LossesShinsaku Sakaue, Han Bao, Yuzhou CaoICML 2026
- Joint Learning of Energy-based Models and their Partition FunctionMichael Eli Sander, Vincent Roulet, Tianlin Liu, Mathieu BlondelICML 2025
Builds on5
- On Gradient Descent Ascent for Nonconvex-Concave Minimax ProblemsTianyi Lin, Chi Jin, Michael I. JordanICML 2020 · 587 citations
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent et al.NeurIPS 2021 · 106 citations
- Super-efficiency of automatic differentiation for functions defined as a minimumPierre Ablin, Gabriel Peyré, Thomas MoreauICML 2020 · 42 citations
- Primal Wasserstein Imitation LearningRobert Dadashi, Léonard Hussenot, Matthieu Geist, Olivier PietquinICLR 2021 · 41 citations
Related papers
- Structured Energy Network As a LossJay Yoon Lee, Dhruvesh Patel, Purujit Goyal, Wenlong Zhao et al.NeurIPS 2022 · 5 citations
- Training Deep Energy-Based Models with f-Divergence MinimizationLantao Yu, Yang Song, Jiaming Song, Stefano ErmonICML 2020 · 50 citations
- Sparse and Structured Hopfield NetworksSaul José Rodrigues dos Santos, Vlad Niculae, Daniel C. McNamee, André F. T. MartinsICML 2024 · 12 citations
- EGC: Image Generation and Classification via a Diffusion Energy-Based ModelQiushan Guo, Chuofan Ma, Yi Jiang, Zehuan Yuan et al.ICCV 2023 · 16 citations
- Adversarial Localized Energy Network for Structured PredictionPingbo Pan, Ping Liu, Yan Yan, Tianbao Yang et al.AAAI 2020 · 9 citations
