LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction
Kanghao Chen, Hangyu Li, Jiazhou Zhou, Zeyu Wang, Lin Wang
摘要
Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard computer vision. However, this task remains challenging due to its inherently ill-posed nature: event cameras only detect the edge and motion information locally. Consequently, the reconstructed videos are often plagued by artifacts and regional blur, primarily caused by the ambiguous semantics of event data. In this paper, we find language naturally conveys abundant semantic information, rendering it stunningly superior in ensuring semantic consistency for E2V reconstruction. Accordingly, we propose a novel framework, called LaSe-E2V, that can achieve semantic-aware high-quality E2V reconstruction from a language-guided perspective, buttressed by the text-conditional diffusion models. However, due to diffusion models' inherent diversity and randomness, it is hardly possible to directly apply them to achieve spatial and temporal consistency for E2V reconstruction. Thus, we first propose an Event-guided Spatiotemporal Attention (ESA) module to condition the event data to the denoising pipeline effectively. We then introduce an event-aware mask loss to ensure temporal coherence and a noise initialization strategy to enhance spatial consistency. Given the absence of event-text-video paired data, we aggregate existing E2V datasets and generate textual descriptions using the tagging models for training and evaluation. Extensive experiments on three datasets covering diverse challenging scenarios (e.g., fast motion, low light) demonstrate the superiority of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion PipelineKanghao Chen, Zixin Zhang, Guoqiang Liang, Lutao Jiang 等NeurIPS 2025 · 被引用 2 次
- EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D ReconstructionKanghao Chen, Zixin Zhang, Hangyu Li, Lin Wang 等AAAI 2026
- LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion ModelsCheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin 等SIGGRAPH 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- From Events to Clarity: The Event-Guided Diffusion Framework for DehazingLing Wang, Yunfan Lu, Wenzong Ma, Huizai Yao 等CVPR 2026
- Event-Diffusion: Event-Based Image Reconstruction and Restoration with Diffusion ModelsQuanmin Liang, Xiawu Zheng, Kai Huang, Yan Zhang 等ACM MM 2023 · 被引用 13 次
- EvIntSR-Net: Event Guided Multiple Latent Frames Reconstruction and Super-resolutionJin Han, Yixin Yang, Chu Zhou, Chao Xu 等ICCV 2021 · 被引用 57 次
- E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep LearningQiang Qu, Yiran Shen, Xiaoming Chen, Yuk Ying Chung 等AAAI 2024 · 被引用 19 次
- NEC-Diff: Noise-Robust Event-RAW Complementary Diffusion for Seeing Motion in Extreme DarknessHaoyue Liu, Jinghan Xu, Luxin Feng, Hanyu Zhou 等CVPR 2026
