RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
Hanqing Liu, Mingjie Liu, Luoping Cui, Endian Lin, Donghong Jiang, Chuang Zhu
摘要
Conventional vision-language models (VLMs) struggle to interpret scenes captured under adverse conditions (e.g., low light, high dynamic range, or fast motion) because standard RGB images degrade in such environments. Event cameras provide a complementary modality: they asynchronously record per-pixel brightness changes with high temporal resolution and wide dynamic range, preserving motion cues where frames fail. We propose RE-VLM, the first dual-stream vision-language model that jointly leverages RGB images and event streams for robust scene understanding across both normal and challenging conditions. RE-VLM employs parallel RGB and event encoders together with a progressive training strategy that aligns heterogeneous visual features with language. To address the scarcity of RGB-Event-Text supervision, we further propose a graph-driven pipeline that converts synchronized RGB-Event streams into verifiable scene graphs, from which we synthesize captions and question–answer (QA) pairs. To develop and evaluate RE-VLM, we construct two datasets: PEOD-Chat, targeting illumination-challenged scenes, and RGBE-Chat, covering diverse scenarios. On captioning and VQA benchmarks, RE-VLM consistently outperforms state-of-the-art RGB-only and event-only models with comparable parameter counts, with particularly large gains under challenging conditions. These results demonstrate the effectiveness of event-augmented VLMs in achieving robust vision-language understanding across a wide range of real-world environments. Code and datasets will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection Under Challenging ConditionsLuoping Cui, Hanqing Liu, Mingjie Liu, Endian Lin 等AAAI 2026 · 被引用 1 次
- Seeing Motion at Nighttime with an Event CameraHaoyue Liu, Shihan Peng, Lin Zhu, Yi Chang 等CVPR 2024
相关 Paper
- EventDrive: Event Cameras for Vision-Language Driving IntelligenceDongyue Lu, Rong Li, Ao Liang, Lingdong Kong 等CVPR 2026 · 被引用 2 次
- Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language ModelsBaoheng Zhang, Jiahui Liu, Gui Zhao, Weizhou Zhang 等CVPR 2026
- EventGPT: Event Stream Understanding with Multimodal Large Language ModelsShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang 等CVPR 2025
- Complementing Event Streams and RGB Frames for Hand Mesh ReconstructionJianping Jiang, Xinyu Zhou, Bingxuan Wang, Xiaoming Deng 等CVPR 2024
- CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic SegmentationRuihao Xia, Chaoqiang Zhao, Meng Zheng, Ziyan Wu 等ICCV 2023 · 被引用 54 次
