Lune

ICML2025Top-tier venue

TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

Xingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer, Mingyu Liu, Hu Cao, Jiajie Zhang, Venkatnarayanan Lakshminarasimhan, Leah Strand, Alois Knoll

2025Year
3Top-tier citations

Abstract

Figure 1 : TUMTraf VideoQA introduces a comprehensive benchmark for video-level traffic scene understanding. Our baseline model, TraffiX-Qwen, is capable of solving multiple tasks, including video QA, spatio-temporal grounding, and referred object captioning, within a unified model. In our approach, the spatio-temporal location of objects is represented as tuples (c, f n, x, y), where c serves as a unique object identifier, f n denotes the normalized frame timestamp, and (x, y) denote the center of the object in the image, normalized with respect to the image dimensions.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 867e603e-d5ba-4846-973a-606bc50201c5

Cited by top-tier papers3

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines