On Training Implicit Models
Zhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang, Zhouchen Lin
Abstract
This paper focuses on training implicit models of infinite layers. Specifically, previous works employ implicit differentiation and solve the exact gradient for the backward propagation. However, is it necessary to compute such an exact but expensive gradient for training? In this work, we propose a novel gradient estimate for implicit models, named phantom gradient, that 1) forgoes the costly computation of the exact gradient; and 2) provides an update direction empirically preferable to the implicit model training. We theoretically analyze the condition under which an ascent direction of the loss landscape could be found, and provide two specific instantiations of the phantom gradient based on the damped unrolling and Neumann series. Experiments on large-scale tasks demonstrate that these lightweight phantom gradients significantly accelerate the backward passes in training implicit models by roughly 1.7 times, and even boost the performance over approaches based on the exact gradient on ImageNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f87b2e1-8a18-46c9-9adf-36b8f478180cCited by top-tier papers47
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Online Training Through Time for Spiking Neural NetworksMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He et al.NeurIPS 2022 · 121 citations
- One-Step Diffusion Distillation via Deep Equilibrium ModelsZhengyang Geng, Ashwini Pokle, J. Zico KolterNeurIPS 2023 · 83 citations
- Object Representations as Fixed Points: Training Iterative Refinement Algorithms with Implicit DifferentiationMichael Chang, Tom Griffiths, Sergey LevineNeurIPS 2022 · 69 citations
Builds on12
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius et al.ICLR 2020 · 341 citations
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Implicit Graph Neural NetworksFangda Gu, Heng Chang, Wenwu Zhu, Somayeh Sojoudi et al.NeurIPS 2020 · 188 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
- Is Attention Better Than Matrix Decomposition?Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li et al.ICLR 2021 · 171 citations
Related papers
- SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit modelsZaccharie Ramzi, Florian Mannel, Shaojie Bai, Jean-Luc Starck et al.ICLR 2022 · 35 citations
- JFB: Jacobian-Free Backpropagation for Implicit NetworksSamy Wu Fung, Howard Heaton, Qiuwei Li, Daniel McKenzie et al.AAAI 2022 · 123 citations
- : Implicit Layers for Implicit RepresentationsZhichun Huang, Shaojie Bai, J. Zico KolterNeurIPS 2021 · 5 citations
- Implicit Stochastic Gradient Descent for Training Physics-Informed Neural NetworksYe Li, Songcan Chen, Sheng-Jun HuangAAAI 2023 · 5 citations
- Efficient Neural Network Training via Forward and Backward Propagation SparsificationXiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao et al.NeurIPS 2021 · 57 citations
