EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned Learning
Yuhan Liu, Linghui Fu, Zhen Yang, Hao Chen, Youfu Li, Yongjian Deng
摘要
Event cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion within the keyframes and ground truth data used for training and inference in these demanding conditions are frequently overlooked. This oversight impedes the perceptual realism and multi-scene generalization capabilities of existing event-based VFI (E-VFI) methods when generating interpolated frames. Motivated by the observation that semantic-perceptual discrepancies between degraded and pristine images are considerably smaller than their image-level differences, we introduce EPA. This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model’s predictions. Specifically, EPA operates in two phases: it first employs a DINO-based perceptual extractor, a customized style adapter, and a reconstruction generator to derive multi-layered, degradation-insensitive semantic-perceptual features ( S ). Second, a novel Bidirectional Event-Guided Alignment (BEGA) module utilizes deformable convolutions to align perceptual features from keyframes to ground truth with inter-frame temporal guidance extracted from event signals. By decoupling the learning process from direct image-level supervision, EPA enhances model robustness against degraded keyframes and unreliable ground truth information. Extensive experiments demonstrate that this approach yields interpolated frames more consistent with human perceptual preferences. Codes are available at https://github.com/yuhan0802/EPA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- One-Shot Flow, Any-Time Frame: A Bidirectional Warping Framework for Event-Based Video Frame InterpolationLinghui Fu, Yuhan Liu, Hao Chen, Zhen Yang 等CVPR 2026
- LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion ModelsCheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin 等SIGGRAPH 2026
它引用的顶会 Paper19
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale FusionStepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig, Stamatios Georgoulis 等CVPR 2022 · 被引用 126 次
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
- LDMVFI: Video Frame Interpolation with Latent Diffusion ModelsDuolikun Danier, Fan Zhang, David BullAAAI 2024 · 被引用 115 次
- Unifying Motion Deblurring and Frame Interpolation with EventsXiang Zhang, Lei YuCVPR 2022 · 被引用 85 次
相关 Paper
- Video Frame Interpolation via Direct Synthesis with the Event-based ReferenceYuhan Liu, Yongjian Deng, Hao Chen, Zhen YangCVPR 2024
- Training Weakly Supervised Video Frame Interpolation with EventsZhiyang Yu, Yu Zhang, Deyuan Liu, Dongqing Zou 等ICCV 2021 · 被引用 45 次
- Event-based Video Frame Interpolation with Cross-Modal Asymmetric Bidirectional Motion FieldsTaewoo Kim, Yujeong Chae, Hyun-Kurl Jang, Kuk-Jin YoonCVPR 2023
- TTA-EVF: Test-Time Adaptation for Event-based Video Frame Interpolation via Reliable Pixel and Sample EstimationHoonhee Cho, Taewoo Kim, Yuhwan Jeong, Kuk-Jin YoonCVPR 2024
- Exploiting Blurry Representations for Event-guided Video Super-ResolutionZeyu Xiao, Xinchao WangAAAI 2026
