RSTT: Real-time Spatial Temporal Transformer for Space-Time Video Super-Resolution
Zhicheng Geng, Luming Liang, Tianyu Ding, Ilya Zharkov
摘要
Space-time video super-resolution (STVSR) is the task of interpolating videos with both Low Frame Rate (LFR) and Low Resolution (LR) to produce High-Frame-Rate (HFR) and also High-Resolution (HR) counterparts. The existing methods based on Convolutional Neural Network (CNN) succeed in achieving visually satisfied results while suffer from slow inference speed due to their heavy architec-tures. We propose to resolve this issue by using a spatial-temporal transformer that naturally incorporates the spa-tial and temporal super resolution modules into a single model. Unlike CNN-based methods, we do not explic-itly use separated building blocks for temporal interpolations and spatial super-resolutions; instead, we only use a single end-to-end transformer architecture. Specifically, a reusable dictionary is built by encoders based on the in-put LFR and LR frames, which is then utilized in the de-coder part to synthesize the HFR and HR frames. compared with the state-of-the-art TMNet [54], our network is 60% smaller (4.5M vs 12.3M parameters) and 80% faster (26.2fps vs 14.3fps on 720 x 576 frames) without sacri-ficing much performance. The source code is available at https://github.com/llmpass/RSTT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan 等NeurIPS 2022 · 被引用 318 次
- Spherical Space Feature Decomposition for Guided Depth Map Super-ResolutionZixiang Zhao, Jiangshe Zhang, Xiang Gu, Chengli Tan 等ICCV 2023 · 被引用 55 次
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-ResolutionYi-Hsin Chen, Si-Cun Chen, Yi-Hsin Chen, Yen-Yu Lin 等ICCV 2023 · 被引用 27 次
- PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementZhanfeng Feng, Long Peng, Xin Di, Yong Guo 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
相关 Paper
- Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-ResolutionXiaoyu Xiang, Yapeng Tian, Yulun Zhang, Yun Fu 等CVPR 2020
- Temporal Modulation Network for Controllable Space-Time Video Super-ResolutionGang Xu, Jun Xu, Zhen Li, Liang Wang 等CVPR 2021
- Space-Time-Aware Multi-Resolution Video EnhancementMuhammad Haris, Greg Shakhnarovich, Norimichi UkitaCVPR 2020
- Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual LearningMengshun Hu, Kui Jiang, Liang Liao, Jing Xiao 等CVPR 2022 · 被引用 37 次
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
