SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events
Li Xu, He Huang, Jun Liu
摘要
Traffic event cognition and reasoning in videos is an important task that has a wide range of applications in intelligent transportation, assisted driving, and autonomous vehicles. In this paper, we create a novel dataset, SUTD-TrafficQA (Traffic Question Answering), which takes the form of video QA based on the collected 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive capability of causal inference and event understanding models in complex traffic scenarios. Specifically, we propose 6 challenging reasoning tasks corresponding to various traffic scenarios, so as to evaluate the reasoning capability over different kinds of complex yet practical traffic events. Moreover, we propose Eclipse, a novel Efficient glimpse network via dynamic inference, in order to achieve computation-efficient and reliable video reasoning. The experiments show that our method achieves superior performance while reducing the computation cost significantly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous DrivingYongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li 等ICLR 2026 · 被引用 196 次
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park 等CVPR 2026 · 被引用 144 次
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds 等NeurIPS 2021 · 被引用 87 次
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- Tem-adapter: Adapting Image-Text Pretraining for Video Question AnswerGuangyi Chen, Xiao Liu, Guangrun Wang, Kun Zhang 等ICCV 2023 · 被引用 32 次
它引用的顶会 Paper7
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 被引用 257 次
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 被引用 173 次
- KnowIT VQA: Answering Knowledge-Based Questions about VideosNoa Garcia, Mayu Otani, Chenhui Chu, Yuta NakashimaAAAI 2020 · 被引用 93 次
- In Defense of Grid Features for Visual Question AnsweringHuaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller 等CVPR 2020
相关 Paper
- A Study of Situational Reasoning for Traffic UnderstandingJiarui Zhang, Filip Ilievski, Kaixin Ma, Aravinda Kollaa 等KDD 2023 · 被引用 8 次
- Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic EnvironmentsDifei Gao, Ruiping Wang, Ziyi Bai, Xilin ChenICCV 2021 · 被引用 36 次
- Visual Traffic Knowledge Graph Generation from Scene ImagesYunfei Guo, Fei Yin, Xiao-Hui Li, Xudong Yan 等ICCV 2023 · 被引用 18 次
- Fine-Grained Evaluation of Large Vision-Language Models in Autonomous DrivingYue Li, Meng Tian, Zhenyu Lin, Jiangtong Zhu 等ICCV 2025 · 被引用 4 次
- STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving ScenesKeishi Ishihara, Kento Sasaki, Tsubasa Takahashi, Daiki Shiono 等AAAI 2026 · 被引用 4 次
