EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned Learning
Yuhan Liu, Linghui Fu, Zhen Yang, Hao Chen, Youfu Li, Yongjian Deng
Abstract
Event cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion within the keyframes and ground truth data used for training and inference in these demanding conditions are frequently overlooked. This oversight impedes the perceptual realism and multi-scene generalization capabilities of existing event-based VFI (E-VFI) methods when generating interpolated frames. Motivated by the observation that semantic-perceptual discrepancies between degraded and pristine images are considerably smaller than their image-level differences, we introduce EPA. This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model’s predictions. Specifically, EPA operates in two phases: it first employs a DINO-based perceptual extractor, a customized style adapter, and a reconstruction generator to derive multi-layered, degradation-insensitive semantic-perceptual features ( S ). Second, a novel Bidirectional Event-Guided Alignment (BEGA) module utilizes deformable convolutions to align perceptual features from keyframes to ground truth with inter-frame temporal guidance extracted from event signals. By decoupling the learning process from direct image-level supervision, EPA enhances model robustness against degraded keyframes and unreliable ground truth information. Extensive experiments demonstrate that this approach yields interpolated frames more consistent with human perceptual preferences. Codes are available at https://github.com/yuhan0802/EPA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- One-Shot Flow, Any-Time Frame: A Bidirectional Warping Framework for Event-Based Video Frame InterpolationLinghui Fu, Yuhan Liu, Hao Chen, Zhen Yang et al.CVPR 2026
- LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion ModelsCheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin et al.SIGGRAPH 2026
Builds on19
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale FusionStepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig, Stamatios Georgoulis et al.CVPR 2022 · 126 citations
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- LDMVFI: Video Frame Interpolation with Latent Diffusion ModelsDuolikun Danier, Fan Zhang, David BullAAAI 2024 · 115 citations
- Unifying Motion Deblurring and Frame Interpolation with EventsXiang Zhang, Lei YuCVPR 2022 · 85 citations
Related papers
- Video Frame Interpolation via Direct Synthesis with the Event-based ReferenceYuhan Liu, Yongjian Deng, Hao Chen, Zhen YangCVPR 2024
- Training Weakly Supervised Video Frame Interpolation with EventsZhiyang Yu, Yu Zhang, Deyuan Liu, Dongqing Zou et al.ICCV 2021 · 45 citations
- Event-based Video Frame Interpolation with Cross-Modal Asymmetric Bidirectional Motion FieldsTaewoo Kim, Yujeong Chae, Hyun-Kurl Jang, Kuk-Jin YoonCVPR 2023
- TTA-EVF: Test-Time Adaptation for Event-based Video Frame Interpolation via Reliable Pixel and Sample EstimationHoonhee Cho, Taewoo Kim, Yuhwan Jeong, Kuk-Jin YoonCVPR 2024
- Exploiting Blurry Representations for Event-guided Video Super-ResolutionZeyu Xiao, Xinchao WangAAAI 2026
