Stabilizing Equilibrium Models by Jacobian Regularization
Shaojie Bai, Vladlen Koltun, J. Zico Kolter
Abstract
Deep equilibrium networks (DEQs) are a new class of models that eschews traditional depth in favor of finding the fixed point of a single nonlinear layer. These models have been shown to achieve performance competitive with the state-of-the-art deep networks while using significantly less memory. Yet they are also slower, brittle to architectural choices, and introduce potential instability to the model. In this paper, we propose a regularization scheme for DEQ models that explicitly regularizes the Jacobian of the fixed-point update equations to stabilize the learning of equilibrium models. We show that this regularization adds only minimal computational cost, significantly stabilizes the fixed-point convergence in both forward and backward passes, and scales well to high-dimensional, realistic domains (e.g., WikiText-103 language modeling and ImageNet classification). Using this method, we demonstrate, for the first time, an implicit-depth model that runs with approximately the same speed and level of performance as popular conventional deep networks such as ResNet-101, while still maintaining the constant memory footprint and architectural simplicity of DEQs. Code is available at https://github.com/locuslab/deq .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- On Training Implicit ModelsZhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang et al.NeurIPS 2021 · 111 citations
- One-Step Diffusion Distillation via Deep Equilibrium ModelsZhengyang Geng, Ashwini Pokle, J. Zico KolterNeurIPS 2023 · 83 citations
- Looped Transformers are Better at Learning Learning AlgorithmsLiu Yang, Kangwook Lee, Robert D. Nowak, Dimitris PapailiopoulosICLR 2024 · 82 citations
- LyaNet: A Lyapunov Framework for Training Neural ODEsIvan Dario Jimenez Rodriguez, Aaron D. Ames, Yisong YueICML 2022 · 78 citations
- Object Representations as Fixed Points: Training Iterative Refinement Algorithms with Implicit DifferentiationMichael Chang, Tom Griffiths, Sergey LevineNeurIPS 2022 · 69 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
Related papers
- Neural Deep Equilibrium SolversShaojie Bai, Vladlen Koltun, J. Zico KolterICLR 2022 · 36 citations
- SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit modelsZaccharie Ramzi, Florian Mannel, Shaojie Bai, Jean-Luc Starck et al.ICLR 2022 · 35 citations
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 7 citations
- Deep Equilibrium Optical Flow EstimationShaojie Bai, Zhengyang Geng, Yash Savani, J. Zico KolterCVPR 2022 · 45 citations
- Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsZhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen LinNeurIPS 2024 · 3 citations
