An In-depth Study of Stochastic Backpropagation
Jun Fang, Mingze Xu, Hao Chen, Bing Shuai, Zhuowen Tu, Joseph Tighe
Abstract
In this paper, we provide an in-depth study of Stochastic Backpropagation (SBP) when training deep neural networks for standard image classification and object detection tasks. During backward propagation, SBP calculates the gradients by only using a subset of feature maps to save the GPU memory and computational cost. We interpret SBP as an efficient way to implement stochastic gradient decent by performing backpropagation dropout, which leads to considerable memory saving and training process speedup, with a minimal impact on the overall model accuracy. We offer some good practices to apply SBP in training image recognition models, which can be adopted in learning a wide range of deep neural networks. Experiments on image classification and object detection show that SBP can save up to 40% of GPU memory with less than 1% accuracy degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 1,747 citations
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
Related papers
- Stochastic Backpropagation: A Memory Efficient Strategy for Training Video ModelsFeng Cheng, Mingze Xu, Yuanjun Xiong, Hao Chen et al.CVPR 2022 · 11 citations
- ReSprop: Reuse Sparsified BackpropagationNegar Goli, Tor M. AamodtCVPR 2020
- DropIT: Dropping Intermediate Tensors for Memory-Efficient DNN TrainingJoya Chen, Kai Xu, Yuhui Wang, Yifei Cheng et al.ICLR 2023 · 2 citations
- DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward PropagationSunghyeon Woo, Baeseong Park, Byeongwook Kim, Minjung Jo et al.NeurIPS 2024 · 13 citations
- Optimal Gradient Checkpoint Search for Arbitrary Computation GraphsJianwei Feng, Dong HuangCVPR 2021
