On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers
Kenji Kawaguchi
摘要
A deep equilibrium model uses implicit layers, which are implicitly defined through an equilibrium point of an infinite sequence of computation. It avoids any explicit computation of the infinite sequence by finding an equilibrium point directly via root-finding and by computing gradients via implicit differentiation. In this paper, we analyze the gradient dynamics of deep equilibrium models with nonlinearity only on weight matrices and non-convex objective functions of weights for regression and classification. Despite non-convexity, convergence to global optimum at a linear rate is guaranteed without any assumption on the width of the models, allowing the width to be smaller than the output dimension and the number of data points. Moreover, we prove a relation between the gradient dynamics of the deep implicit layer and the dynamics of trust region Newton method of a shallow explicit layer. This mathematically proven relation along with our numerical observation suggests the importance of understanding implicit bias of implicit layers and an open problem on the topic. Our proofs deal with implicit layers, weight tying and nonlinearity on weights, and differ from those in the related literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- On Training Implicit ModelsZhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang 等NeurIPS 2021 · 被引用 111 次
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 被引用 87 次
- Deep Equilibrium Approaches to Diffusion ModelsAshwini Pokle, Zhengyang Geng, J. Zico KolterNeurIPS 2022 · 被引用 61 次
- Robust Implicit Networks via Non-Euclidean ContractionsSaber Jafarpour, Alexander Davydov, Anton V. Proskurnikov, Francesco BulloNeurIPS 2021 · 被引用 59 次
- EIGNN: Efficient Infinite-Depth Graph Neural NetworksJuncheng Liu, Kenji Kawaguchi, Bryan Hooi, Yiwei Wang 等NeurIPS 2021 · 被引用 56 次
它引用的顶会 Paper3
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 被引用 44 次
相关 Paper
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 被引用 7 次
- Understanding the Dynamics of Gradient Flow in Overparameterized Linear modelsSalma Tarmoun, Guilherme França, Benjamin D. Haeffele, René VidalICML 2021 · 被引用 76 次
- A global convergence theory for deep ReLU implicit networks via over-parameterizationTianxiang Gao, Hailiang Liu, Jia Liu, Hridesh Rajan 等ICLR 2022 · 被引用 21 次
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 被引用 80 次
- Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsZhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen LinNeurIPS 2024 · 被引用 3 次
