Efficient and Information-Preserving Future Frame Prediction and Beyond
Wei Yu, Yichao Lu, Steve Easterbrook, Sanja Fidler
Abstract
Applying resolution-preserving blocks is a common practice to maximize information preservation in video prediction, yet their high memory consumption greatly limits their application scenarios. We propose CrevNet, a Conditionally Reversible Network that uses reversible architectures to build a bijective two-way autoencoder and its complementary recurrent predictor. Our model enjoys the theoretically guaranteed property of no information loss during the feature extraction, much lower memory consumption and computational efficiency. The lightweight nature of our model enables us to incorporate 3D convolutions without concern of memory bottleneck, enhancing the model's ability to capture both short-term and long-term temporal dependencies. Our proposed approach achieves state-of-the-art results on Moving MNIST, Traffic4cast and KITTI datasets. We further demonstrate the transferability of our self-supervised learning method by exploiting its learnt features for object detection on KITTI. Our competitive results indicate the potential of using CrevNet as a generative pre-training strategy to guide downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0c8d0cf-b39d-4c7b-adad-019b0b7bb8caCited by top-tier papers22
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 313 citations
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.NeurIPS 2021 · 193 citations
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 117 citations
- STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video PredictionZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.CVPR 2022 · 57 citations
- MMVP: Motion-Matrix-based Video PredictionYiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich NeumannICCV 2023 · 39 citations
Related papers
- Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient FinetuningChen Zhao, Shuming Liu, Karttikeya Mangalam, Guocheng Qian et al.CVPR 2024
- Anomaly Detection in Video via Self-Supervised and Multi-Task LearningMariana-Iuliana Georgescu, Antonio Barbalau, Radu Tudor Ionescu, Fahad Shahbaz Khan et al.CVPR 2021
- Reversible Vision TransformersKarttikeya Mangalam, Haoqi Fan, Yanghao Li, Chao-Yuan Wu et al.CVPR 2022 · 52 citations
- Self-Supervised Learning of Compressed Video RepresentationsYoungjae Yu, Sangho Lee, Gunhee Kim, Yale SongICLR 2021 · 15 citations
- Does Visual Pretraining Help End-to-End Reasoning?Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab et al.NeurIPS 2023 · 4 citations
