Stochastic Backpropagation: A Memory Efficient Strategy for Training Video Models
Feng Cheng, Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Li, Wei Xia
Abstract
We propose a memory efficient method, named Stochastic Backpropagation (SBP), for training deep neural networks on videos. It is based on the finding that gradients from incomplete execution for backpropagation can still effectively train the models with minimal accuracy loss, which attributes to the high redundancy of video. SBP keeps all forward paths but randomly and independently removes the backward paths for each network layer in each training step. It reduces the GPU memory cost by eliminating the need to cache activation values corresponding to the dropped backward paths, whose amount can be controlled by an adjustable keep-ratio. Experiments show that SBP can be applied to a wide range of models for video tasks, leading to up to 80.0% GPU memory saving and 10% training speedup with less than 1% accuracy drop on action recognition and temporal action detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a9dd485-859e-4ab9-84ce-367db724d647Cited by top-tier papers8
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 35 citations
- To Adapt or Not to Adapt? Real-Time Adaptation for Semantic SegmentationMarc Botet Colomer, Pier Luigi Dovesi, Theodoros Panagiotakopoulos, Joao Frederico Carvalho et al.ICCV 2023 · 19 citations
- E2E-LOAD: End-to-End Long-form Online Action DetectionShuqiang Cao, Weixin Luo, Bairui Wang, Wei Zhang et al.ICCV 2023 · 12 citations
- Uncovering the Unseen: Discover Hidden Intentions by Micro-Behavior Graph ReasoningZhuo Zhou, Wenxuan Liu, Danni Xu, Zheng Wang et al.ACM MM 2023 · 9 citations
- An In-depth Study of Stochastic BackpropagationJun Fang, Mingze Xu, Hao Chen, Bing Shuai et al.NeurIPS 2022 · 2 citations
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
Related papers
- Sideways: Depth-Parallel Training of Video ModelsMateusz Malinowski, Grzegorz Swirszcz, João Carreira, Viorica PatrauceanCVPR 2020
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
- ReSprop: Reuse Sparsified BackpropagationNegar Goli, Tor M. AamodtCVPR 2020
- Gradient Forward-Propagation for Large-Scale Temporal Video ModellingMateusz Malinowski, Dimitrios Vytiniotis, Grzegorz Swirszcz, Viorica Patraucean et al.CVPR 2021
- DropIT: Dropping Intermediate Tensors for Memory-Efficient DNN TrainingJoya Chen, Kai Xu, Yuhui Wang, Yifei Cheng et al.ICLR 2023 · 2 citations
