SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit models
Zaccharie Ramzi, Florian Mannel, Shaojie Bai, Jean-Luc Starck, Philippe Ciuciu, Thomas Moreau
Abstract
In recent years, implicit deep learning has emerged as a method to increase the effective depth of deep neural networks. While their training is memory-efficient, they are still significantly slower to train than their explicit counterparts. In Deep Equilibrium Models (DEQs), the training is performed as a bi-level problem, and its computational complexity is partially driven by the iterative inversion of a huge Jacobian matrix. In this paper, we propose a novel strategy to tackle this computational bottleneck from which many bi-level problems suffer. The main idea is to use the quasi-Newton matrices from the forward pass to efficiently approximate the inverse Jacobian matrix in the direction needed for the gradient computation. We provide a theorem that motivates using our method with the original forward algorithms. In addition, by modifying these forward algorithms, we further provide theoretical guarantees that our method asymptotically estimates the true implicit gradient. We empirically study this approach and the recent Jacobian-Free method in different settings, ranging from hyperparameter optimization to large Multiscale DEQs (MDEQs) applied to CIFAR and ImageNet. Both methods reduce significantly the computational cost of the backward pass. While SHINE has a clear advantage on hyperparameter optimization problems, both methods attain similar computational performances for larger scale problems such as MDEQs at the cost of a limited performance drop compared to the original models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63e5231d-adf3-4f1c-8099-95b20e1d2070Cited by top-tier papers11
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- A framework for bilevel optimization that enables stochastic and global variance reduction algorithmsMathieu Dagréou, Pierre Ablin, Samuel Vaiter, Thomas MoreauNeurIPS 2022 · 149 citations
- MGNNI: Multiscale Graph Neural Networks with Implicit LayersJuncheng Liu, Bryan Hooi, Kenji Kawaguchi, Xiaokui XiaoNeurIPS 2022 · 36 citations
- One-step differentiation of iterative algorithmsJérôme Bolte, Edouard Pauwels, Samuel VaiterNeurIPS 2023 · 36 citations
- A Closer Look at the Adversarial Robustness of Deep Equilibrium ModelsZonghan Yang, Tianyu Pang, Yang LiuNeurIPS 2022 · 20 citations
Builds on2
Related papers
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 80 citations
- Neural Deep Equilibrium SolversShaojie Bai, Vladlen Koltun, J. Zico KolterICLR 2022 · 36 citations
- JFB: Jacobian-Free Backpropagation for Implicit NetworksSamy Wu Fung, Howard Heaton, Qiuwei Li, Daniel McKenzie et al.AAAI 2022 · 123 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
- Deep Equilibrium Optical Flow EstimationShaojie Bai, Zhengyang Geng, Yash Savani, J. Zico KolterCVPR 2022 · 45 citations
