Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics
Zhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen Lin
Abstract
The deep equilibrium model (DEQ) generalizes the conventional feedforward 1 neural network by fixing the same weights for each layer block and extending 2 the number of layers to infinity. This novel model directly finds the fixed points 3 of such a forward process as features for prediction. Despite empirical evidence 4 showcasing its efficacy compared to feedforward neural networks, a theoretical 5 understanding for its separation and bias is still limited. In this paper, we take a 6 step by proposing some separations and studying the bias of DEQ in its expressive 7 power and learning dynamics. The results include: (1) A general separation is 8 proposed, showing the existence of a width-m DEQ that any fully connected neural 9 networks (FNNs) with depth O ( m α ) for α ∈ (0 , 1) cannot approximate unless 10 its width is sub-exponential in m ; (2) DEQ with polynomially bounded size and 11 magnitude can efficiently approximate certain steep functions (which has very large 12 derivatives) in L ∞ norm, whereas FNN with bounded depth and exponentially 13 bounded width cannot unless its weights magnitudes are exponentially large; (3) 14 The implicit regularization caused by gradient flow from a diagonal linear DEQ 15 is characterized, with specific examples showing the benefits brought by such 16 regularization. From the overall study, a high-level conjecture from our analysis 17 and empirical validations is that DEQ has potential advantages in learning certain 18 high-frequency components. 19
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5c1eaa3-6e53-4f1f-be84-790c017e44a8Cited by top-tier papers2
- Expressive Power of Implicit Models: Rich Equilibria and Test-Time ScalingJialin Liu, Lisang Ding, Stanley J. Osher, Wotao YinICLR 2026 · 3 citations
- Consistency Deep Equilibrium ModelsJunchao Lin, Zenan Ling, Jingwen Xu, Robert QiuICML 2026 · 2 citations
Builds on18
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee et al.NeurIPS 2020 · 98 citations
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.ICML 2020 · 83 citations
- Stabilizing Equilibrium Models by Jacobian RegularizationShaojie Bai, Vladlen Koltun, J. Zico KolterICML 2021 · 80 citations
Related papers
- Wide Neural Networks as Gaussian Processes: Lessons from Deep Equilibrium ModelsTianxiang Gao, Xiaokai Huo, Hailiang Liu, Hongyang GaoNeurIPS 2023 · 20 citations
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 47 citations
- Neural Deep Equilibrium SolversShaojie Bai, Vladlen Koltun, J. Zico KolterICLR 2022 · 36 citations
- Positive Concave Deep Equilibrium ModelsMateusz Gabor, Tomasz Piotrowski, Renato L. G. CavalcanteICML 2024 · 7 citations
- GEQ: Gaussian Kernel Inspired Equilibrium ModelsMingjie Li, Yisen Wang, Zhouchen LinNeurIPS 2023 · 1 citation
