TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
Xingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer, Mingyu Liu, Hu Cao, Jiajie Zhang, Venkatnarayanan Lakshminarasimhan, Leah Strand, Alois Knoll
Abstract
Figure 1 : TUMTraf VideoQA introduces a comprehensive benchmark for video-level traffic scene understanding. Our baseline model, TraffiX-Qwen, is capable of solving multiple tasks, including video QA, spatio-temporal grounding, and referred object captioning, within a unified model. In our approach, the spatio-temporal location of objects is represented as tuples (c, f n, x, y), where c serves as a unique object identifier, f n denotes the normalized frame timestamp, and (x, y) denote the center of the object in the image, normalized with respect to the image dimensions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 867e603e-d5ba-4846-973a-606bc50201c5Cited by top-tier papers3
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving ScenesKeishi Ishihara, Kento Sasaki, Tsubasa Takahashi, Daiki Shiono et al.AAAI 2026 · 4 citations
- TSBOW - Traffic Surveillance Benchmark for Occluded Vehicles Under Various Weather ConditionsNgoc Doan-Minh Huynh, Duong Nguyen-Ngoc Tran, Long Hoang Pham, Tai Huu-Phuong Tran et al.AAAI 2026
Builds on20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen et al.NeurIPS 2024 · 6,113 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic VideosFanheng Kong, Jingyuan Zhang, Hongzhi Zhang, Shi Feng et al.ACL 2025
- VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and UnderstandingShibo Gao, Peipei Yang, Yangyang Liu, Yi Chen et al.AAAI 2026 · 5 citations
- SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language ModelsYe Sun, Hao Zhang, Henghui Ding, Tiehua Zhang et al.NeurIPS 2025 · 9 citations
- VideoLoom: A Video Large Language Model for Joint Spatial-Temporal UnderstandingJiapeng Shi, junke Wang, Zuyao You, Bo He et al.ICML 2026 · 5 citations
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He et al.CVPR 2024 · 18 citations
