One Forward is Enough for Neural Network Training via Likelihood Ratio Method
Jinyang Jiang, Zeliang Zhang, Chenliang Xu, Zhaofei Yu, Yijie Peng
摘要
While backpropagation (BP) is the mainstream approach for gradient computation in neural network training, its heavy reliance on the chain rule of differentiation constrains the designing flexibility of network architecture and training pipelines. We avoid the recursive computation in BP and develop a unified likelihood ratio (ULR) method for gradient estimation with just one forward propagation. Not only can ULR be extended to train a wide variety of neural network architectures, but the computation flow in BP can also be rearranged by ULR for better device adaptation. Moreover, we propose several variance reduction techniques to further accelerate the training process. Our experiments offer numerical results across diverse aspects, including various neural network training scenarios, computation flow rearrangement, and fine-tuning of pre-trained models. All findings demonstrate that ULR effectively enhances the flexibility of neural network training by permitting localized module training without compromising the global objective and significantly boosts the network robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning to Transform Dynamically for Better Adversarial TransferabilityRongyi Zhu, Zeliang Zhang, Zhuo Liu, Chenliang Xu 等CVPR 2024 · 被引用 18 次
- Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio OptimizerTao Ren, Zishi Zhang, Jinyang Jiang, Zehao Li 等ICLR 2026 · 被引用 5 次
- Discover and Mitigate Multiple Biased Subgroups in Image ClassifiersZeliang Zhang, Mingqian Feng, Zhiheng Li, Chenliang XuCVPR 2024 · 被引用 5 次
- Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural NetworksMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He 等ICLR 2026 · 被引用 2 次
- FLOPS: Forward Learning with OPtimal SamplingTao Ren, Zishi Zhang, Jinyang Jiang, Guanghao Li 等ICLR 2025
它引用的顶会 Paper3
- The HSIC Bottleneck: Deep Learning without Back-PropagationKurt Wan-Duo Ma, J. P. Lewis, W. Bastiaan KleijnAAAI 2020 · 被引用 180 次
- Training Spiking Neural Networks with Event-driven BackpropagationYaoyu Zhu, Zhaofei Yu, Wei Fang, Xiaodong Xie 等NeurIPS 2022 · 被引用 57 次
- TACR-Net: Editing on Deep Video and Voice PortraitsLuchuan Song, Bin Liu, Guojun Yin, Xiaoyi Dong 等ACM MM 2021 · 被引用 20 次
相关 Paper
- Efficient Neural Network Training via Forward and Backward Propagation SparsificationXiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao 等NeurIPS 2021 · 被引用 57 次
- Accelerated training through iterative gradient propagation along the residual pathErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICLR 2025
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 被引用 3 次
- Backpropagation-Free Deep Learning with Recursive Local Representation AlignmentAlexander G. Ororbia II, Ankur Mali, Daniel Kifer, C. Lee GilesAAAI 2023 · 被引用 19 次
- Efficient Backpropagation with Variance Controlled Adaptive SamplingZiteng Wang, Jianfei Chen, Jun ZhuICLR 2024 · 被引用 5 次
