RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
Chirag Parikh, Deepti Rawat, Rakshitha R. T, Tathagata Ghosh, Ravi Kiran Sarvadevabhatla
摘要
We introduce RoadSocial, a large-scale, diverse VideoQA dataset tailored for generic road event understanding from social media narratives. Unlike existing datasets limited by regional bias, viewpoint bias and expert-driven annotations, RoadSocial captures the global complexity of road events with varied geographies, camera viewpoints (CCTV, handheld, drones) and rich social discourse. Our scalable semi-automatic annotation framework leverages Text LLMs and Video LLMs to generate comprehensive questionanswer pairs across 12 challenging QA tasks, pushing the boundaries of road event understanding. RoadSocial is derived from social media videos spanning 14M frames and 414K social comments, resulting in a dataset with 13.2K videos, 674 tags and 260K high-quality QA pairs. We evaluate 18 Video LLMs (open-source and proprietary, drivingspecific and general-purpose) on our road event understanding benchmark. We also demonstrate RoadSocial's utility in improving road event understanding capabilities of general-purpose Video LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao 等AAAI 2024 · 被引用 314 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su 等CVPR 2024
- Abductive Ego-View Accident Video Understanding for Safe Driving PerceptionJianwu Fang, Lei-Lei Li, Junfei Zhou, Junbin Xiao 等CVPR 2024
相关 Paper
- NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language ModelsSung-Yeon Park, Can Cui, Yunsheng Ma, Ahmadreza Moradipari 等ICCV 2025 · 被引用 6 次
- STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving ScenesKeishi Ishihara, Kento Sasaki, Tsubasa Takahashi, Daiki Shiono 等AAAI 2026 · 被引用 4 次
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingJingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang 等NeurIPS 2025 · 被引用 25 次
- Artemis: Towards Referential Understanding in Complex VideosJihao Qiu, Yuan Zhang, Xi Tang, Lingxi Xie 等NeurIPS 2024 · 被引用 33 次
