Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics
Zhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen Lin
摘要
The deep equilibrium model (DEQ) generalizes the conventional feedforward 1 neural network by fixing the same weights for each layer block and extending 2 the number of layers to infinity. This novel model directly finds the fixed points 3 of such a forward process as features for prediction. Despite empirical evidence 4 showcasing its efficacy compared to feedforward neural networks, a theoretical 5 understanding for its separation and bias is still limited. In this paper, we take a 6 step by proposing some separations and studying the bias of DEQ in its expressive 7 power and learning dynamics. The results include: (1) A general separation is 8 proposed, showing the existence of a width-m DEQ that any fully connected neural 9 networks (FNNs) with depth O ( m α ) for α ∈ (0 , 1) cannot approximate unless 10 its width is sub-exponential in m ; (2) DEQ with polynomially bounded size and 11 magnitude can efficiently approximate certain steep functions (which has very large 12 derivatives) in L ∞ norm, whereas FNN with bounded depth and exponentially 13 bounded width cannot unless its weights magnitudes are exponentially large; (3) 14 The implicit regularization caused by gradient flow from a diagonal linear DEQ 15 is characterized, with specific examples showing the benefits brought by such 16 regularization. From the overall study, a high-level conjecture from our analysis 17 and empirical validations is that DEQ has potential advantages in learning certain 18 high-frequency components. 19
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Expressive Power of Implicit Models: Rich Equilibria and Test-Time ScalingJialin Liu, Lisang Ding, Stanley J. Osher, Wotao YinICLR 2026 · 被引用 3 次
- Consistency Deep Equilibrium ModelsJunchao Lin, Zenan Ling, Jingwen Xu, Robert QiuICML 2026 · 被引用 2 次
它引用的顶会 Paper18
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 被引用 177 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等ICML 2020 · 被引用 83 次
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 被引用 80 次
相关 Paper
- Wide Neural Networks as Gaussian Processes: Lessons from Deep Equilibrium ModelsTianxiang Gao, Xiaokai Huo, Hailiang Liu, Hongyang GaoNeurIPS 2023 · 被引用 20 次
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 被引用 47 次
- Neural Deep Equilibrium SolversShaojie Bai, Vladlen Koltun, J. Zico KolterICLR 2022 · 被引用 36 次
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 被引用 7 次
- GEQ: Gaussian Kernel Inspired Equilibrium ModelsMingjie Li, Yisen Wang, Zhouchen LinNeurIPS 2023 · 被引用 1 次
