RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
Hanqing Liu, Mingjie Liu, Luoping Cui, Endian Lin, Donghong Jiang, Chuang Zhu
Abstract
Conventional vision-language models (VLMs) struggle to interpret scenes captured under adverse conditions (e.g., low light, high dynamic range, or fast motion) because standard RGB images degrade in such environments. Event cameras provide a complementary modality: they asynchronously record per-pixel brightness changes with high temporal resolution and wide dynamic range, preserving motion cues where frames fail. We propose RE-VLM, the first dual-stream vision-language model that jointly leverages RGB images and event streams for robust scene understanding across both normal and challenging conditions. RE-VLM employs parallel RGB and event encoders together with a progressive training strategy that aligns heterogeneous visual features with language. To address the scarcity of RGB-Event-Text supervision, we further propose a graph-driven pipeline that converts synchronized RGB-Event streams into verifiable scene graphs, from which we synthesize captions and question–answer (QA) pairs. To develop and evaluate RE-VLM, we construct two datasets: PEOD-Chat, targeting illumination-challenged scenes, and RGBE-Chat, covering diverse scenarios. On captioning and VQA benchmarks, RE-VLM consistently outperforms state-of-the-art RGB-only and event-only models with comparable parameter counts, with particularly large gains under challenging conditions. These results demonstrate the effectiveness of event-augmented VLMs in achieving robust vision-language understanding across a wide range of real-world environments. Code and datasets will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection Under Challenging ConditionsLuoping Cui, Hanqing Liu, Mingjie Liu, Endian Lin et al.AAAI 2026 · 1 citation
- Seeing Motion at Nighttime with an Event CameraHaoyue Liu, Shihan Peng, Lin Zhu, Yi Chang et al.CVPR 2024
Related papers
- EventDrive: Event Cameras for Vision-Language Driving IntelligenceDongyue Lu, Rong Li, Ao Liang, Lingdong Kong et al.CVPR 2026 · 2 citations
- Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language ModelsBaoheng Zhang, Jiahui Liu, Gui Zhao, Weizhou Zhang et al.CVPR 2026
- EventGPT: Event Stream Understanding with Multimodal Large Language ModelsShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang et al.CVPR 2025
- Complementing Event Streams and RGB Frames for Hand Mesh ReconstructionJianping Jiang, Xinyu Zhou, Bingxuan Wang, Xiaoming Deng et al.CVPR 2024
- CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic SegmentationRuihao Xia, Chaoqiang Zhao, Meng Zheng, Ziyan Wu et al.ICCV 2023 · 54 citations
