Bregman Neural Networks
Jordan Frécon, Gilles Gasso, Massimiliano Pontil, Saverio Salzo
Abstract
We present a framework based on bilevel optimization for learning multilayer, deep data representations. On the one hand, the lower-level problem finds a representation by successively minimizing layer-wise objectives made of the sum of a prescribed regularizer, a fidelity term and a linear function depending on the representation found at the previous layer. On the other hand, the upper-level problem optimizes over the linear functions to yield a linearly separable final representation. We show that, by choosing the fidelity term as the quadratic distance between two successive layer-wise representations, the bilevel problem reduces to the training of a feedforward neural network. Instead, by elaborating on Bregman distances, we devise a novel neural network architecture additionally involving the inverse of the activation function reminiscent of the skip connection used in ResNets. Numerical experiments suggest that the proposed Bregman variant benefits from better learning properties and more robust prediction performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfc85ff2-96c4-4b67-9f84-98a5370d8972Cited by top-tier papers4
- Transformers from an Optimization PerspectiveYongyi Yang, Zengfeng Huang, David P. WipfNeurIPS 2022 · 44 citations
- Attention as Implicit Structural InferenceRyan Singh, Christopher L. BuckleyNeurIPS 2023 · 12 citations
- A Constrained Optimization Perspective of Unrolled TransformersJavier Porras-Valenzuela, Samar Hadou, Alejandro RibeiroICML 2026
- A Bregman Proximal Viewpoint on Neural OperatorsAbdel-Rahim Mezidi, Jordan Patracone, Saverio Salzo, Amaury Habrard et al.ICML 2025
Builds on1
Related papers
- Enhanced Bilevel Optimization via Bregman DistanceFeihu Huang, Junyi Li, Shangqian Gao, Heng HuangNeurIPS 2022 · 41 citations
- Functional Bilevel Optimization for Machine LearningIeva Petrulionyte, Julien Mairal, Michael ArbelNeurIPS 2024 · 27 citations
- Overcoming Lower-Level Constraints in Bilevel Optimization: A Novel Approach with Regularized Gap FunctionsWei Yao, Haian Yin, Shangzhi Zeng, Jin ZhangICLR 2025
- Neur2BiLO: Neural Bilevel OptimizationJustin Dumouchelle, Esther Julien, Jannis Kurtz, Elias B. KhalilNeurIPS 2024 · 10 citations
- Towards the Training of Deeper Predictive Coding Neural NetworksChang Qi, Matteo Forasassi, Thomas Lukasiewicz, Tommaso SalvatoriICML 2026 · 6 citations
