A Multigrid Method for Efficiently Training Video Models
Chao-Yuan Wu, Ross B. Girshick, Kaiming He, Christoph Feichtenhofer, Philipp Krähenbühl
Abstract
Training competitive deep video models is an order of magnitude slower than training their counterpart image models. Slow training causes long research cycles, which hinders progress in video understanding research. Following standard practice for training image models, video model training has used a fixed mini-batch shape: a specific number of clips, frames, and spatial size. However, what is the optimal shape? High resolution models perform well, but train slowly. Low resolution models train faster, but are less accurate. Inspired by multigrid methods in numerical optimization, we propose to use variable mini-batch shapes with different spatial-temporal resolutions that are varied according to a schedule. The different shapes arise from resampling the training data on multiple sampling grids. Training is accelerated by scaling up the mini-batch size and learning rate when shrinking the other dimensions. We empirically demonstrate a general and robust grid schedule that yields a significant out-of-the-box training speedup without a loss in accuracy for different models (I3D, non-local, SlowFast), datasets (Kinetics, Something-Something, Charades), and training settings (with and without pre-training, 128 GPUs or 1 GPU). As an illustrative example, the proposed multigrid method trains a ResNet-50 SlowFast network 4.5x faster (wall-clock time, same hardware) while also improving accuracy (+0.8% absolute) on Kinetics-400 compared to baseline training. Code is available online.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57aeacb1-9887-4146-8fd5-543faaa43207Cited by top-tier papers23
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek et al.NeurIPS 2023 · 303 citations
- Refining activation downsampling with SoftPoolAlexandros Stergiou, Ronald Poppe, Grigorios KalliatakisICCV 2021 · 195 citations
- Space-time Mixing Attention for Video TransformerAdrian Bulat, Juan-Manuel Pérez-Rúa, Swathikiran Sudhakaran, Brais Martínez et al.NeurIPS 2021 · 158 citations
- DirecFormer: A Directed Attention in Transformer Approach to Robust Action RecognitionThanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo et al.CVPR 2022 · 70 citations
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
Related papers
- Accelerating the Training of Video Super-resolution ModelsLijian Lin, Xintao Wang, Zhongang Qi, Ying ShanAAAI 2023 · 4 citations
- PGT: A Progressive Method for Training Models on Long VideosBo Pang, Gao Peng, Yizhuo Li, Cewu LuCVPR 2021
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- X3D: Expanding Architectures for Efficient Video RecognitionChristoph FeichtenhoferCVPR 2020
- DSANet: Dynamic Segment Aggregation Network for Video-Level Representation LearningWenhao Wu, Yuxiang Zhao, Yanwu Xu, Xiao Tan et al.ACM MM 2021 · 30 citations
