On Training Implicit Models
Zhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang, Zhouchen Lin
摘要
This paper focuses on training implicit models of infinite layers. Specifically, previous works employ implicit differentiation and solve the exact gradient for the backward propagation. However, is it necessary to compute such an exact but expensive gradient for training? In this work, we propose a novel gradient estimate for implicit models, named phantom gradient, that 1) forgoes the costly computation of the exact gradient; and 2) provides an update direction empirically preferable to the implicit model training. We theoretically analyze the condition under which an ascent direction of the loss landscape could be found, and provide two specific instantiations of the phantom gradient based on the damped unrolling and Neumann series. Experiments on large-scale tasks demonstrate that these lightweight phantom gradients significantly accelerate the backward passes in training implicit models by roughly 1.7 times, and even boost the performance over approaches based on the exact gradient on ImageNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- Online Training Through Time for Spiking Neural NetworksMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He 等NeurIPS 2022 · 被引用 121 次
- One-Step Diffusion Distillation via Deep Equilibrium ModelsZhengyang Geng, Ashwini Pokle, J. Zico KolterNeurIPS 2023 · 被引用 83 次
- Object Representations as Fixed Points: Training Iterative Refinement Algorithms with Implicit DifferentiationMichael Chang, Tom Griffiths, Sergey LevineNeurIPS 2022 · 被引用 69 次
它引用的顶会 Paper12
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius 等ICLR 2020 · 被引用 341 次
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Implicit Graph Neural NetworksFangda Gu, Heng Chang, Wenwu Zhu, Somayeh Sojoudi 等NeurIPS 2020 · 被引用 188 次
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 被引用 177 次
- Is Attention Better Than Matrix Decomposition?Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li 等ICLR 2021 · 被引用 171 次
相关 Paper
- SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit modelsZaccharie Ramzi, Florian Mannel, Shaojie Bai, Jean-Luc Starck 等ICLR 2022 · 被引用 35 次
- JFB: Jacobian-Free Backpropagation for Implicit NetworksSamy Wu Fung, Howard Heaton, Qiuwei Li, Daniel McKenzie 等AAAI 2022 · 被引用 123 次
- : Implicit Layers for Implicit RepresentationsZhichun Huang, Shaojie Bai, J. Zico KolterNeurIPS 2021 · 被引用 5 次
- Implicit Stochastic Gradient Descent for Training Physics-Informed Neural NetworksYe Li, Songcan Chen, Sheng-Jun HuangAAAI 2023 · 被引用 5 次
- Efficient Neural Network Training via Forward and Backward Propagation SparsificationXiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao 等NeurIPS 2021 · 被引用 57 次
