Deep equilibrium networks are sensitive to initialization statistics
Atish Agarwala, Samuel S. Schoenholz
2022年份
12被引次数
6顶会引用
摘要
Deep equilibrium networks (DEQs) are a promising way to construct models which trade off memory for compute. However, theoretical understanding of these models is still lacking compared to traditional networks, in part because of the repeated application of a single set of weights. We show that DEQs are sensitive to the higher order statistics of the matrix families from which they are initialized. In particular, initializing with orthogonal or symmetric matrices allows for greater stability in training. This gives us a practical prescription for initializations which allow for training with a broader range of initial weight scales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Exploiting Connections between Lipschitz Structures for Certifiably Robust Deep Equilibrium ModelsAaron J. Havens, Alexandre Araujo, Siddharth Garg, Farshad Khorrami 等NeurIPS 2023 · 被引用 15 次
- Towards training digitally-tied analog blocks via hybrid gradient computationTimothy Nest, Maxence ErnoultNeurIPS 2024 · 被引用 6 次
- Deep Equilibrium Models are Almost Equivalent to Not-so-deep Explicit Models for High-dimensional Gaussian MixturesZenan Ling, Longbo Li, Zhanbo Feng, Yixuan Zhang 等ICML 2024 · 被引用 6 次
- Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsZhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen LinNeurIPS 2024 · 被引用 3 次
- Derivative Informed Learning of Exchange-Correlation FunctionalsEike S. Eberhard, Luca Anthony Thiede, Abdulrahman Aldossary, Andreas Burger 等ICML 2026
它引用的顶会 Paper8
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita 等NeurIPS 2020 · 被引用 261 次
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 被引用 177 次
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 被引用 169 次
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear NetworksWei Hu, Lechao Xiao, Jeffrey PenningtonICLR 2020 · 被引用 136 次
相关 Paper
- Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?Yaniv Blumenfeld, Dar Gilboa, Daniel SoudryICML 2020 · 被引用 18 次
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 被引用 80 次
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 被引用 7 次
- Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationCem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein 等NeurIPS 2022 · 被引用 43 次
- Lyapunov-Stable Deep Equilibrium ModelsHaoyu Chu, Shikui Wei, Ting Liu, Yao Zhao 等AAAI 2024 · 被引用 10 次
