Video Frame Prediction from a Single Image and Events
Juanjuan Zhu, Zhexiong Wan, Yuchao Dai
Abstract
Recently, the task of Video Frame Prediction (VFP), which predicts future video frames from previous ones through extrapolation, has made remarkable progress. However, the performance of existing VFP methods is still far from satisfactory due to the fixed framerate video used: 1) they have difficulties in handling complex dynamic scenes; 2) they cannot predict future frames with flexible prediction time intervals. The event cameras can record the intensity changes asynchronously with a very high temporal resolution, which provides rich dynamic information about the observed scenes. In this paper, we propose to predict video frames from a single image and the following events, which can not only handle complex dynamic scenes but also predict future frames with flexible prediction time intervals. First, we introduce a symmetrical cross-modal attention augmentation module to enhance the complementary information between images and events. Second, we propose to jointly achieve optical flow estimation and frame generation by combining the motion information of events and the semantic information of the image, then inpainting the holes produced by forward warping to obtain an ideal prediction frame. Based on these, we propose a lightweight pyramidal coarse-to-fine model that can predict a 720P frame within 25 ms. Extensive experiments show that our proposed model significantly outperforms the state-of-the-art frame-based and event-based VFP methods and has the fastest runtime. Code is available at https://npucvr.github.io/VFPSIE/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70e7df41-c3a2-4301-9716-4183bf938c2eCited by top-tier papers1
Ask how each one uses itBuilds on21
- Channel Attention Is All You Need for Video Frame InterpolationMyungsub Choi, Heewon Kim, Bohyung Han, Ning Xu et al.AAAI 2020 · 362 citations
- A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionZhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 341 citations
- IFRNet: Intermediate Feature Refine Network for Efficient Frame InterpolationLingtong Kong, Boyuan Jiang, Donghao Luo, Wenqing Chu et al.CVPR 2022 · 166 citations
- Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale FusionStepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig, Stamatios Georgoulis et al.CVPR 2022 · 126 citations
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
Related papers
- Event-based Video Frame Interpolation with Cross-Modal Asymmetric Bidirectional Motion FieldsTaewoo Kim, Yujeong Chae, Hyun-Kurl Jang, Kuk-Jin YoonCVPR 2023
- Training Weakly Supervised Video Frame Interpolation with EventsZhiyang Yu, Yu Zhang, Deyuan Liu, Dongqing Zou et al.ICCV 2021 · 45 citations
- Video Frame Interpolation via Direct Synthesis with the Event-based ReferenceYuhan Liu, Yongjian Deng, Hao Chen, Zhen YangCVPR 2024
- E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic ScenesYan Liu, Zehao Chen, Haojie Yan, De Ma et al.ICCV 2025 · 2 citations
- Frequency-Aware Event-Based Video Deblurring for Real-World Motion BlurTaewoo Kim, Hoonhee Cho, Kuk-Jin YoonCVPR 2024
