DFDNet: Disentangling and Filtering Dynamics for Enhanced Video Prediction
Lianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao, Yi Bin
Abstract
Videos inherently contain complex temporal dynamics across various spatial directions, often entangled in ways that obscure effective dynamic extraction. Previous studies typically process video spatiotemporal features without disentangling, which hampers their ability to extract dynamic information. Additionally, the extraction of dynamics is disrupted by transient high-dynamic information in video sequences, e.g., noise or flicker, which has received limited attention in the literature. To tackle those problems, this paper proposes the Disentangling and Filtering Dynamics Network (DFD-Net). Firstly, to disentangle the interwoven dynamics, DFD-Net decomposes the spatially encoded video sequences into lower dimensional sequences. Secondly, a learnable threshold filter is proposed to eliminate the transient high-dynamic information. Thirdly, the model incorporates an MLP to extract the temporal dependencies from the disentangled and filtered sequences. DFDNet demonstrates competitive performance across four chosen datasets, including both low and highresolution videos. Specifically, on the low-resolution Moving MNIST dataset, DFDNet achieves a 19% improvement on MSE over the previous state-of-the-art model. On the highresolution SJTU4K dataset, it outperforms the previous stateof-the-art model by 10% on the LPIPS metric under similar inference time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d537e666-0159-4bcc-a2cb-c9ca10e0efdbCited by top-tier papers1
Ask how each one uses itBuilds on13
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 313 citations
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 127 citations
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 117 citations
- STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video PredictionZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.CVPR 2022 · 57 citations
- MMVP: Motion-Matrix-based Video PredictionYiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich NeumannICCV 2023 · 39 citations
Related papers
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou et al.NeurIPS 2023 · 27 citations
- FastDVDnet: Towards Real-Time Deep Video Denoising Without Flow EstimationMatias Tassano, Julie Delon, Thomas VeitCVPR 2020
- Temporal Denoising Mask Synthesis Network for Learning Blind Video Temporal ConsistencyYifeng Zhou, Xing Xu, Fumin Shen, Lianli Gao et al.ACM MM 2020 · 9 citations
- Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer LearningZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yingya Zhang et al.ICCV 2023 · 40 citations
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu et al.AAAI 2024 · 13 citations
