Relational Space-Time Query in Long-Form Videos
Xitong Yang, Fu-Jen Chu, Matt Feiszli, Raghav Goyal, Lorenzo Torresani, Du Tran
Abstract
Q: When did I do activity 𝑎 that involves interaction with fefobject 𝑜 ? A: Temporal locations of the corresponding activity fef [𝑠!, 𝑒!] !"# % Figure 1. Illustration of the three types of queries in our Relational Space-Time Query (ReST) framework. Given a long video spanning up to 30 minutes, a set of queries are provided to assess a model's ability to understand activities, objects, and their interactions in the video. All queries and answers are generated in the form of pre-defined templates (top-left) to avoid the ambiguity and bias introduced by language input / output. Note that ReST is a holistic framework that supports constructing queries with different levels of complexity beyond the three basic types described in this paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03daf167-de98-48d8-8fa3-8e6c776e24e8Cited by top-tier papers4
- A Simple LLM Framework for Long-Range Video Question-AnsweringCe Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang et al.EMNLP 2024 · 37 citations
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 34 citations
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 5 citations
- DrVideo: Document Retrieval Based Long Video UnderstandingZiyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun et al.CVPR 2025
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Mitigating Query Selection Bias in Referring Video Object SegmentationDingwei Zhang, Dong Zhang, Jinhui TangACM MM 2025 · 1 citation
- Evaluating Temporal Queries Over Video FeedsYueting Chen, Xiaohui Yu, Nick Koudas, Ziqiang YuSIGMOD 2021 · 17 citations
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee et al.CVPR 2020
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesXingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer et al.ICML 2025
