You Only Align Once: Bidirectional Interaction for Spatial-Temporal Video Super-Resolution
Mengshun Hu, Kui Jiang, Zhixiang Nie, Zheng Wang
Abstract
Spatial-Temporal Video Super-Resolution (ST-VSR) technology generates high-quality videos with higher resolution and higher frame rates. Existing advanced methods accomplish ST-VSR tasks through the association of Spatial and Temporal video super-resolution (S-VSR and T-VSR). These methods require two alignments and fusions in S-VSR and T-VSR, which is obviously redundant and fails to sufficiently explore the information flow of consecutive spatial LR frames. Although bidirectional learning (future-to-past and past-to-future) was introduced to cover all input frames, the direct fusion of final predictions fails to sufficiently exploit intrinsic correlations of bidirectional motion learning and spatial information from all frames. We propose an effective yet efficient recurrent network with bidirectional interaction for ST-VSR, where only one alignment and fusion is needed. Specifically, it first performs backward inference from future to past, and then follows forward inference to super-resolve intermediate frames. The backward and forward inferences are assigned to learn structures and details to simplify the learning task with joint optimizations. Furthermore, a Hybrid Fusion Module (HFM) is designed to aggregate and distill information to refine spatial information and reconstruct high-quality video frames. Extensive experiments on two public datasets demonstrate that our method outperforms state-of-the-art methods in efficiency, and reduces calculation cost by about 22%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b778e7f-64dd-480b-b51f-5cbb3204d2a2Cited by top-tier papers3
- Store and Fetch Immediately: Everything Is All You Need for Space-Time Video Super-resolutionMengshun Hu, Kui Jiang, Zhixiang Nie, Jiahuan Zhou et al.AAAI 2023 · 13 citations
- MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-ResolutionHua Chang, Xin Xu, Wei Liu, Wei Wang et al.AAAI 2026 · 1 citation
- Time Without Time: Pseudo-Temporal Representation for Space-Time Super-ResolutionHee Min Choi, Hyoa Kang, Suji Kim, Dokwan Oh et al.CVPR 2026
Builds on17
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang et al.ICCV 2019 · 309 citations
- XVFI: eXtreme Video Frame InterpolationHyeonjun Sim, Jihyong Oh, Munchurl KimICCV 2021 · 207 citations
- Asymmetric Bilateral Motion Estimation for Video Frame InterpolationJunheum Park, Chul Lee, Chang-Su KimICCV 2021 · 186 citations
- Video Frame Interpolation via Deformable Separable ConvolutionXianhang Cheng, Zhenzhong ChenAAAI 2020 · 153 citations
Related papers
- Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual LearningMengshun Hu, Kui Jiang, Liang Liao, Jing Xiao et al.CVPR 2022 · 37 citations
- How Video Super-Resolution and Frame Interpolation Mutually BenefitChengcheng Zhou, Zongqing Lu, Linge Li, Qiangyu Yan et al.ACM MM 2021 · 12 citations
- Stereo Video Super-Resolution via Exploiting View-Temporal CorrelationsRuikang Xu, Zeyu Xiao, Mingde Yao, Yueyi Zhang et al.ACM MM 2021 · 20 citations
- Temporal Modulation Network for Controllable Space-Time Video Super-ResolutionGang Xu, Jun Xu, Zhen Li, Liang Wang et al.CVPR 2021
- Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-ResolutionXiaoyu Xiang, Yapeng Tian, Yulun Zhang, Yun Fu et al.CVPR 2020
