Event-based Video Reconstruction Using Transformer
Wenming Weng, Yueyi Zhang, Zhiwei Xiong
Abstract
Event cameras, which output events by detecting spatio- temporal brightness changes, bring a novel paradigm to image sensors with high dynamic range and low latency. Previous works have achieved impressive performances on event-based video reconstruction by introducing convolutional neural networks (CNNs). However, intrinsic locality of convolutional operations is not capable of modeling long-range dependency, which is crucial to many vision tasks. In this paper, we present a hybrid CNN- Transformer network for event-based video reconstruction (ET-Net), which merits the fine local information from CNN and global contexts from Transformer In addition, we further propose a Token Pyramid Aggregation strategy to implement multi-scale token integration for relating internal and intersected semantic concepts in the token-space. Experimental results demonstrate that our proposed method achieves superior performance over state-of-the-art methods on multiple real-world event datasets. The code is available at https://github.com/WarranWeng/ET-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9297f0de-ce61-47f5-b32a-b540042a8ed4Cited by top-tier papers49
- Event-based Video Reconstruction via Potential-assisted Spiking Neural NetworkLin Zhu, Xiao Wang, Yi Chang, Jianing Li et al.CVPR 2022 · 109 citations
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun et al.ICCV 2023 · 86 citations
- Coherent Event Guided Low-Light Video EnhancementJinxiu Liang, Yixin Yang, Boyu Li, Peiqi Duan et al.ICCV 2023 · 54 citations
- Label-Free Event-based Object Recognition via Joint Learning with Image Reconstruction from EventsHoonhee Cho, Hyeonseong Kim, Yujeong Chae, Kuk-Jin YoonICCV 2023 · 38 citations
- Person Re-Identification without Identification via Event AnonymizationShafiq Ahmad, Pietro Morerio, Alessio Del BueICCV 2023 · 36 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Space-Time Video Super-Resolution Using Temporal ProfilesZeyu Xiao, Zhiwei Xiong, Xueyang Fu, Dong Liu et al.ACM MM 2020 · 54 citations
- Stereo Video Super-Resolution via Exploiting View-Temporal CorrelationsRuikang Xu, Zeyu Xiao, Mingde Yao, Yueyi Zhang et al.ACM MM 2021 · 20 citations
- Back to Event Basics: Self-Supervised Learning of Image Reconstruction for Event Cameras via Photometric ConstancyFederico Paredes-Vallés, Guido C. H. E. de CroonCVPR 2021
- Space-Time Distillation for Video Super-ResolutionZeyu Xiao, Xueyang Fu, Jie Huang, Zhen Cheng et al.CVPR 2021
Related papers
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang et al.NeurIPS 2025 · 4 citations
- EvIntSR-Net: Event Guided Multiple Latent Frames Reconstruction and Super-resolutionJin Han, Yixin Yang, Chu Zhou, Chao Xu et al.ICCV 2021 · 57 citations
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
- Learning To Reconstruct High Speed and High Dynamic Range Videos From EventsYunhao Zou, Yinqiang Zheng, Tsuyoshi Takatani, Ying FuCVPR 2021
- Learning to Super Resolve Intensity Images From EventsS. Mohammad Mostafavi I., Jonghyun Choi, Kuk-Jin YoonCVPR 2020
