EventDrive: Event Cameras for Vision-Language Driving Intelligence
Dongyue Lu, Rong Li, Ao Liang, Lingdong Kong, Wei Yin, Lai Xing Ng, Benoit R. Cottereau, Camille Simon Chane, Wei Tsang Ooi
Abstract
Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame-based sensors and capturing temporal structure that conventional exposures often miss. These properties make events a powerful complement to RGB in autonomous driving, especially under blur, glare, and rapid motion, where frame-based perception can become unreliable. However, existing event-aware vision-language models remain limited to generic perception and do not reveal how event sensing contributes to reasoning and decision-making across the full driving loop. We present EventDrive, a large-scale benchmark and model suite that unifies event streams, RGB frames, and language supervision across four core dimensions: Perception, Understanding, Prediction, and Planning, covering captions, structured QA, grounding, motion-state recognition, trajectory forecasting, and planning tasks. Building on this foundation, EventDrive-VLM introduces a multi-horizon event pyramid and a temporal-horizon mixture-of-experts module to adaptively encode and fuse asynchronous and frame-based information for downstream reasoning. Comprehensive evaluation across diverse tasks shows that event streams provide substantial gains in temporal precision, motion awareness, and robustness, bringing event sensing into the center of driving intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c6d08a6-1096-4969-bea2-c2e5ec586e0cBuilds on27
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci et al.NeurIPS 2020 · 381 citations
Related papers
- RE-VLM: Event-Augmented Vision-Language Model for Scene UnderstandingHanqing Liu, Mingjie Liu, Luoping Cui, Endian Lin et al.CVPR 2026
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li et al.NeurIPS 2025 · 7 citations
- Ev-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event CamerasHoonhee Cho, Jae-Young Kang, Youngho Kim, Kuk-Jin YoonCVPR 2025
- When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid NetworkDong Xiao, Guangyao Chen, Peixi Peng, Yangru Huang et al.ICML 2025
- Unsupervised 3d Motion Estimation Using Event CameraHan Han, Wei Zhai, Tiesong Zhao, Bin Li et al.CVPR 2026
