Accelerating the Training of Video Super-resolution Models
Lijian Lin, Xintao Wang, Zhongang Qi, Ying Shan
摘要
Despite that convolution neural networks (CNN) have recently demonstrated high-quality reconstruction for video super-resolution (VSR), efficiently training competitive VSR models remains a challenging problem. It usually takes an order of magnitude more time than training their counterpart image models, leading to long research cycles. Existing VSR methods typically train models with fixed spatial and temporal sizes from beginning to end. The fixed sizes are usually set to large values for good performance, resulting to slow training. However, is such a rigid training strategy necessary for VSR? In this work, we show that it is possible to gradually train video models from small to large spatial/temporal sizes, i.e., in an easy-to-hard manner. In particular, the whole training is divided into several stages and the earlier stage has smaller training spatial shape. Inside each stage, the temporal size also varies from short to long while the spatial size remains unchanged. Training is accelerated by such a multigrid training strategy, as most of computation is performed on smaller spatial and shorter temporal shapes. For further acceleration with GPU parallelization, we also investigate the large minibatch training without the loss in accuracy. Extensive experiments demonstrate that our method is capable of largely speeding up training (up to 6.2× speedup in wall-clock training time) without performance drop for various VSR models. The code is available at https://github.com/TencentARC/Efficient-VSR-Training .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 被引用 522 次
- Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang 等ICCV 2019 · 被引用 309 次
- Understanding Deformable Alignment in Video Super-ResolutionKelvin C. K. Chan, Xintao Wang, Ke Yu, Chao Dong 等AAAI 2021 · 被引用 184 次
- TDAN: Temporally-Deformable Alignment Network for Video Super-ResolutionYapeng Tian, Yulun Zhang, Yun Fu, Chenliang XuCVPR 2020
相关 Paper
- A Multigrid Method for Efficiently Training Video ModelsChao-Yuan Wu, Ross B. Girshick, Kaiming He, Christoph Feichtenhofer 等CVPR 2020
- SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame MaskZekun Ai, Xiaotong Luo, Yanyun Qu, Yuan XieACM MM 2024 · 被引用 2 次
- Parallel Training of GRU Networks with a Multi-Grid Solver for Long SequencesEuhyun Moon, Eric C. CyrICLR 2022 · 被引用 9 次
- RSTT: Real-time Spatial Temporal Transformer for Space-Time Video Super-ResolutionZhicheng Geng, Luming Liang, Tianyu Ding, Ilya ZharkovCVPR 2022 · 被引用 103 次
- Time Without Time: Pseudo-Temporal Representation for Space-Time Super-ResolutionHee Min Choi, Hyoa Kang, Suji Kim, Dokwan Oh 等CVPR 2026
