Continuous Space-Time Video Super-Resolution with 3D Fourier Fields
Alexander Becker, Julius Erbach, Dominik Narnhofer, Konrad Schindler
Abstract
We introduce a novel formulation for continuous space-time video super-resolution. Instead of decoupling the representation of a video sequence into separate spatial and temporal components and relying on brittle, explicit frame warping for motion compensation, we encode video as a continuous, spatio-temporally coherent 3D Video Fourier Field (VFF). That representation offers three key advantages: (1) it enables cheap, flexible sampling at arbitrary locations in space and time; (2) it is able to simultaneously capture fine spatial detail and smooth temporal dynamics; and (3) it offers the possibility to include an analytical, Gaussian point spread function in the sampling to ensure aliasing-free reconstruction at arbitrary scale. The coefficients of the proposed, Fourier-like sinusoidal basis are predicted with a neural encoder with a large spatio-temporal receptive field, conditioned on the low-resolution input video. Through extensive experiments, we show that our joint modeling substantially improves both spatial and temporal super-resolution and sets a new state of the art for multiple benchmarks: across a wide range of upscaling factors, it delivers sharper and temporally more consistent reconstructions than existing baselines, while being computationally more efficient. Project page: https://v3vsr.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan et al.NeurIPS 2022 · 318 citations
- Local Texture Estimator for Implicit Representation FunctionJaewon Lee, Kyong Hwan JinCVPR 2022 · 193 citations
- Learning Trajectory-Aware Transformer for Video Super-ResolutionChengxu Liu, Huan Yang, Jianlong Fu, Xueming QianCVPR 2022 · 113 citations
- VideoINR: Learning Video Implicit Neural Representation for Continuous Space-Time Super-ResolutionZeyuan Chen, Yinbo Chen, Jingwen Liu, Xingqian Xu et al.CVPR 2022 · 95 citations
- MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-ResolutionYi-Hsin Chen, Si-Cun Chen, Yi-Hsin Chen, Yen-Yu Lin et al.ICCV 2023 · 27 citations
Related papers
- BF-STVSR: B-Splines and Fourier - Best Friends for High Fidelity Spatial-Temporal Video Super-ResolutionEunjin Kim, Hyeonjin Kim, Kyong Hwan Jin, Jaejun YooCVPR 2025
- Bias for Action: Video Implicit Neural Representations with Bias ModulationAlper Kayabasi, Anil Kumar Vadathya, Guha Balakrishnan, Vishwanath SaragadamCVPR 2025
- Enhancing Video Super-Resolution via Implicit Resampling-based AlignmentKai Xu, Ziwei Yu, Xin Wang, Michael Bi Mi et al.CVPR 2024 · 22 citations
- GaussianVideo: Efficient Video Representation via Hierarchical Gaussian SplattingAndrew Bond, Jui-Hsien Wang, Long Mai, Erkut Erdem et al.ICCV 2025 · 14 citations
- Arbitrary-Scale Video Super-resolution Guided by Dynamic ContextCong Huang, Jiahao Li, Lei Chu, Dong Liu et al.AAAI 2024 · 5 citations
