DFDNet: Disentangling and Filtering Dynamics for Enhanced Video Prediction
Lianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao, Yi Bin
摘要
Videos inherently contain complex temporal dynamics across various spatial directions, often entangled in ways that obscure effective dynamic extraction. Previous studies typically process video spatiotemporal features without disentangling, which hampers their ability to extract dynamic information. Additionally, the extraction of dynamics is disrupted by transient high-dynamic information in video sequences, e.g., noise or flicker, which has received limited attention in the literature. To tackle those problems, this paper proposes the Disentangling and Filtering Dynamics Network (DFD-Net). Firstly, to disentangle the interwoven dynamics, DFD-Net decomposes the spatially encoded video sequences into lower dimensional sequences. Secondly, a learnable threshold filter is proposed to eliminate the transient high-dynamic information. Thirdly, the model incorporates an MLP to extract the temporal dependencies from the disentangled and filtered sequences. DFDNet demonstrates competitive performance across four chosen datasets, including both low and highresolution videos. Specifically, on the low-resolution Moving MNIST dataset, DFDNet achieves a 19% improvement on MSE over the previous state-of-the-art model. On the highresolution SJTU4K dataset, it outperforms the previous stateof-the-art model by 10% on the LPIPS metric under similar inference time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 被引用 313 次
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 被引用 127 次
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 被引用 117 次
- STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video PredictionZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 等CVPR 2022 · 被引用 57 次
- MMVP: Motion-Matrix-based Video PredictionYiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich NeumannICCV 2023 · 被引用 39 次
相关 Paper
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou 等NeurIPS 2023 · 被引用 27 次
- FastDVDnet: Towards Real-Time Deep Video Denoising Without Flow EstimationMatias Tassano, Julie Delon, Thomas VeitCVPR 2020
- Temporal Denoising Mask Synthesis Network for Learning Blind Video Temporal ConsistencyYifeng Zhou, Xing Xu, Fumin Shen, Lianli Gao 等ACM MM 2020 · 被引用 9 次
- Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer LearningZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yingya Zhang 等ICCV 2023 · 被引用 40 次
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu 等AAAI 2024 · 被引用 13 次
