Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction
Beibei Jin, Yu Hu, Qiankun Tang, Jingyu Niu, Zhiping Shi, Yinhe Han, Xiaowei Li
Abstract
Video prediction is a pixel-wise dense prediction task to infer future frames based on past frames. Missing appearance details and motion blur are still two major problems for current models, leading to image distortion and temporal inconsistency. We point out the necessity of exploring multi-frequency analysis to deal with the two problems. Inspired by the frequency band decomposition characteristic of Human Vision System (HVS), we propose a video prediction network based on multi-level wavelet analysis to uniformly deal with spatial and temporal information. Specifically, multi-level spatial discrete wavelet transform decomposes each video frame into anisotropic sub-bands with multiple frequencies, helping to enrich structural information and reserve fine details. On the other hand, multilevel temporal discrete wavelet transform which operates on time axis decomposes the frame sequence into sub-band groups of different frequencies to accurately capture multifrequency motions under a fixed frame rate. Extensive experiments on diverse datasets demonstrate that our model shows significant improvements on fidelity and temporal consistency over the state-of-the-art works. Source code and videos are available at https://github.com/ Bei-Jin/STMFANet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5997cea7-b7ea-4fc8-a898-57ea57761468Cited by top-tier papers19
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 434 citations
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 313 citations
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.NeurIPS 2021 · 193 citations
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 117 citations
Builds on2
Related papers
- WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelZongjian Li, Bin Lin, Yang Ye, Liuhan Chen et al.CVPR 2025
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
- Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time VariationsXuesong Nie, Yunfeng Yan, Siyuan Li, Cheng Tan et al.AAAI 2024 · 31 citations
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
