Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
Lingdong Kong, Dongyue Lu, Alan Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit Cottereau
摘要
Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven object grounding in event-based perception. Built from real-world driving data, we provide over 30,000 validated referring expressions, each enriched with four grounding attributes -- appearance, status, relation to viewer, and relation to other objects -- bridging spatial, temporal, and relational reasoning. To fully exploit these cues, we propose EventRefer, an attribute-aware grounding framework that dynamically fuses multi-attribute representations through a Mixture of Event-Attribute Experts (MoEE). Our method adapts to different modalities and scene dynamics, achieving consistent gains over state-of-the-art baselines in event-only, frame-only, and event-frame fusion settings. We hope our dataset and approach will establish a foundation for advancing multimodal, temporally-aware, and language-driven perception in real-world robotics and autonomy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong 等ICLR 2026 · 被引用 28 次
- EventDrive: Event Cameras for Vision-Language Driving IntelligenceDongyue Lu, Rong Li, Ao Liang, Lingdong Kong 等CVPR 2026 · 被引用 2 次
- ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation PoliciesYiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang 等ICLR 2026
它引用的顶会 Paper50
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 被引用 482 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 被引用 427 次
- Learning to Detect Objects with a 1 Megapixel Event CameraEtienne Perot, Pierre de Tournemire, Davide Nitti, Jonathan Masci 等NeurIPS 2020 · 被引用 381 次
相关 Paper
- LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsHanyu Zhou, Gim Hee LeeICCV 2025 · 被引用 6 次
- Event6D: Event-based Novel Object 6D Pose TrackingJae-Young Kang, Hoonhee Cho, Taeyeop Lee, Minjun Kang 等CVPR 2026 · 被引用 4 次
- ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreJiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, Lin WangCVPR 2024 · 被引用 15 次
- VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMYuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng 等CVPR 2025
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
