Neural Deep Equilibrium Solvers
Shaojie Bai, Vladlen Koltun, J. Zico Kolter
Abstract
A deep equilibrium (DEQ) model abandons traditional depth by solving for the fixed point of a single nonlinear layer f θ . This structure enables decoupling the internal structure of the layer (which controls representational capacity) from how the fixed point is actually computed (which impacts inference-time efficiency), which is usually via classic techniques such as Broyden's method or Anderson acceleration. In this paper, we show that one can exploit such decoupling and substantially enhance this fixed point computation using a custom neural solver. Specifically, our solver uses a parameterized network to both guess an initial value of the optimization and perform iterative updates, in a method that generalizes a learnable form of Anderson acceleration and can be trained end-to-end in an unsupervised manner. Such a solution is particularly well suited to the implicit model setting, because inference in these models requires repeatedly solving for a fixed point of the same nonlinear layer for different inputs, a task at which our network excels. Our experiments show that these neural equilibrium solvers are fast to train (only taking an extra 0.9-1.1% over the original DEQ's training time), require few additional parameters (1-3% of the original model size), yet lead to a 2× speedup in DEQ network inference without any degradation in accuracy across numerous domains and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1c04464-3261-43a6-ad9d-d471a437c101Cited by top-tier papers13
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- Looped Transformers are Better at Learning Learning AlgorithmsLiu Yang, Kangwook Lee, Robert D. Nowak, Dimitris PapailiopoulosICLR 2024 · 82 citations
- Deep Equilibrium Approaches to Diffusion ModelsAshwini Pokle, Zhengyang Geng, J. Zico KolterNeurIPS 2022 · 61 citations
- Towards Scalable and Stable Parallelization of Nonlinear RNNsXavier Gonzalez, Andrew Warrington, Jimmy T. H. Smith, Scott W. LindermanNeurIPS 2024 · 47 citations
- Deep Equilibrium Optical Flow EstimationShaojie Bai, Zhengyang Geng, Yash Savani, J. Zico KolterCVPR 2022 · 45 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Implicit Graph Neural NetworksFangda Gu, Heng Chang, Wenwu Zhu, Somayeh Sojoudi et al.NeurIPS 2020 · 188 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
Related papers
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 80 citations
- Joint inference and input optimization in equilibrium networksSwaminathan Gurumurthy, Shaojie Bai, Zachary Manchester, J. Zico KolterNeurIPS 2021 · 22 citations
- Consistency Deep Equilibrium ModelsJunchao Lin, Zenan Ling, Jingwen Xu, Robert QiuICML 2026 · 2 citations
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 7 citations
- Declarative nets that are equilibrium modelsRussell Tsuchida, Suk Yee Yong, Mohammad Ali Armin, Lars Petersson et al.ICLR 2022 · 5 citations
