V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
Hanyue Lou, Jinxiu Liang, Minggui Teng, Yi Wang, Boxin Shi
Abstract
Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150 times reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours--an order of magnitude larger than existing event datasets, yielding substantial improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8826e5fc-22fc-4853-b234-750d0d8649d4Cited by top-tier papers2
- AE2VID: Event-based Video Reconstruction via Aperture ModulationChenxu Bai, Boyu Li, Peiqi Duan, Xinyu Zhou et al.CVPR 2026 · 1 citation
- Texvent: Asynchronous Event Data Simulation via Text PromptRuofei Wang, Peiqi Duan, Ka Chun Cheung, Simon See et al.CVPR 2026
Builds on10
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 139 citations
- Event-based Video Reconstruction via Potential-assisted Spiking Neural NetworkLin Zhu, Xiao Wang, Yi Chang, Jianing Li et al.CVPR 2022 · 109 citations
- Spatio-Temporal Recurrent Networks for Event-Based Optical Flow EstimationZiluo Ding, Rui Zhao, Jiyuan Zhang, Tianxiao Gao et al.AAAI 2022 · 76 citations
- Learning Optical Flow from Event Camera with Rendered DatasetXinglong Luo, Kunming Luo, Ao Luo, Zhengning Wang et al.ICCV 2023 · 28 citations
Related papers
- Video to Events: Recycling Video Datasets for Event CamerasDaniel Gehrig, Mathias Gehrig, Javier Hidalgo-Carrió, Davide ScaramuzzaCVPR 2020
- How to Learn a Domain-Adaptive Event Simulator?Daxin Gu, Jia Li, Yu Zhang, Yonghong TianACM MM 2021 · 8 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- E2PNet: Event to Point Cloud Registration with Spatio-Temporal Representation LearningXiuhong Lin, Changjie Qiu, Zhipeng Cai, Siqi Shen et al.NeurIPS 2023 · 18 citations
- A Voxel Graph CNN for Object Classification with Event CamerasYongjian Deng, Hao Chen, Hai Liu, Youfu LiCVPR 2022 · 55 citations
