Deep equilibrium networks are sensitive to initialization statistics
Atish Agarwala, Samuel S. Schoenholz
Abstract
Deep equilibrium networks (DEQs) are a promising way to construct models which trade off memory for compute. However, theoretical understanding of these models is still lacking compared to traditional networks, in part because of the repeated application of a single set of weights. We show that DEQs are sensitive to the higher order statistics of the matrix families from which they are initialized. In particular, initializing with orthogonal or symmetric matrices allows for greater stability in training. This gives us a practical prescription for initializations which allow for training with a broader range of initial weight scales.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a9dfccb-7bfe-41ef-9191-7b1f6ccdeac5Cited by top-tier papers6
- Exploiting Connections between Lipschitz Structures for Certifiably Robust Deep Equilibrium ModelsAaron J. Havens, Alexandre Araujo, Siddharth Garg, Farshad Khorrami et al.NeurIPS 2023 · 15 citations
- Towards training digitally-tied analog blocks via hybrid gradient computationTimothy Nest, Maxence ErnoultNeurIPS 2024 · 6 citations
- Deep Equilibrium Models are Almost Equivalent to Not-so-deep Explicit Models for High-dimensional Gaussian MixturesZenan Ling, Longbo Li, Zhanbo Feng, Yixuan Zhang et al.ICML 2024 · 6 citations
- Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsZhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen LinNeurIPS 2024 · 3 citations
- Derivative Informed Learning of Exchange-Correlation FunctionalsEike S. Eberhard, Luca Anthony Thiede, Abdulrahman Aldossary, Andreas Burger et al.ICML 2026
Builds on8
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita et al.NeurIPS 2020 · 261 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear NetworksWei Hu, Lechao Xiao, Jeffrey PenningtonICLR 2020 · 136 citations
Related papers
- Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?Yaniv Blumenfeld, Dar Gilboa, Daniel SoudryICML 2020 · 18 citations
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 80 citations
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 7 citations
- Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationCem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein et al.NeurIPS 2022 · 43 citations
- Lyapunov-Stable Deep Equilibrium ModelsHaoyu Chu, Shikui Wei, Ting Liu, Yao Zhao et al.AAAI 2024 · 10 citations
