TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
Xingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer, Mingyu Liu, Hu Cao, Jiajie Zhang, Venkatnarayanan Lakshminarasimhan, Leah Strand, Alois Knoll
摘要
Figure 1 : TUMTraf VideoQA introduces a comprehensive benchmark for video-level traffic scene understanding. Our baseline model, TraffiX-Qwen, is capable of solving multiple tasks, including video QA, spatio-temporal grounding, and referred object captioning, within a unified model. In our approach, the spatio-temporal location of objects is represented as tuples (c, f n, x, y), where c serves as a unique object identifier, f n denotes the normalized frame timestamp, and (x, y) denote the center of the object in the image, normalized with respect to the image dimensions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma 等AAAI 2026 · 被引用 119 次
- STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving ScenesKeishi Ishihara, Kento Sasaki, Tsubasa Takahashi, Daiki Shiono 等AAAI 2026 · 被引用 4 次
- TSBOW - Traffic Surveillance Benchmark for Occluded Vehicles Under Various Weather ConditionsNgoc Doan-Minh Huynh, Duong Nguyen-Ngoc Tran, Long Hoang Pham, Tai Huu-Phuong Tran 等AAAI 2026
它引用的顶会 Paper20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen 等NeurIPS 2024 · 被引用 6,113 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
相关 Paper
- TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic VideosFanheng Kong, Jingyuan Zhang, Hongzhi Zhang, Shi Feng 等ACL 2025
- VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and UnderstandingShibo Gao, Peipei Yang, Yangyang Liu, Yi Chen 等AAAI 2026 · 被引用 5 次
- SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language ModelsYe Sun, Hao Zhang, Henghui Ding, Tiehua Zhang 等NeurIPS 2025 · 被引用 9 次
- VideoLoom: A Video Large Language Model for Joint Spatial-Temporal UnderstandingJiapeng Shi, junke Wang, Zuyao You, Bo He 等ICML 2026 · 被引用 5 次
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He 等CVPR 2024 · 被引用 18 次
