Relational Space-Time Query in Long-Form Videos
Xitong Yang, Fu-Jen Chu, Matt Feiszli, Raghav Goyal, Lorenzo Torresani, Du Tran
摘要
Q: When did I do activity 𝑎 that involves interaction with fefobject 𝑜 ? A: Temporal locations of the corresponding activity fef [𝑠!, 𝑒!] !"# % Figure 1. Illustration of the three types of queries in our Relational Space-Time Query (ReST) framework. Given a long video spanning up to 30 minutes, a set of queries are provided to assess a model's ability to understand activities, objects, and their interactions in the video. All queries and answers are generated in the form of pre-defined templates (top-left) to avoid the ambiguity and bias introduced by language input / output. Note that ReST is a holistic framework that supports constructing queries with different levels of complexity beyond the three basic types described in this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Simple LLM Framework for Long-Range Video Question-AnsweringCe Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang 等EMNLP 2024 · 被引用 37 次
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 被引用 34 次
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 被引用 5 次
- DrVideo: Document Retrieval Based Long Video UnderstandingZiyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 等CVPR 2025
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Mitigating Query Selection Bias in Referring Video Object SegmentationDingwei Zhang, Dong Zhang, Jinhui TangACM MM 2025 · 被引用 1 次
- Evaluating Temporal Queries Over Video FeedsYueting Chen, Xiaohui Yu, Nick Koudas, Ziqiang YuSIGMOD 2021 · 被引用 17 次
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 等CVPR 2020
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesXingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer 等ICML 2025
