Evaluating Temporal Queries Over Video Feeds
Yueting Chen, Xiaohui Yu, Nick Koudas, Ziqiang Yu
Abstract
Recent advances in Computer Vision and Deep Learning have made possible the efficient extraction of structured information from frames of video feeds. As such, a stream of objects and their associated classes along with unique object identifiers derived via object tracking can be generated, providing unique objects as they are captured across frames. In this paper we initiate a study of temporal queries involving objects and their co-occurrences in video feeds. For example, queries that identify video segments during which the same two red cars and the same two humans appear jointly for five minutes are of interest to many applications ranging from law enforcement to security and safety. We take the first step and define such queries in a way that they incorporate certain physical aspects of video capture such as object occlusion. We present an architecture consisting of three layers, namely object detection/tracking, intermediate data generation, and query evaluation. We propose two techniques, Marked Frame Set (MFS) and Sparse State Graph (SSG), to organize all detected objects in the intermediate data generation layer, which effectively, given the queries, minimizes the number of objects and frames that have to be considered during query evaluation. We also introduce an algorithm called SSG-CM that processes incoming frames against the SSG and efficiently prunes objects and frames unrelated to query evaluation, while maintaining all states required for succinct query evaluation. We present the results of a thorough experimental evaluation utilizing both real and synthetic data, establishing the trade-offs between MFS and SSG. We stress various parameters of interest in our evaluation and demonstrate that the proposed query evaluation methodology coupled with the proposed algorithms is capable to evaluate temporal queries over video feeds efficiently, achieving orders of magnitude performance benefits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f51b9352-453b-45cb-bd40-75c98d115faaCited by top-tier papers7
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes et al.VLDB 2023 · 18 citations
- Spatial and Temporal Constrained Ranked Retrieval over VideosYueting Chen, Nick Koudas, Xiaohui Yu, Ziqiang YuVLDB 2022 · 16 citations
- Accelerating Aggregation Queries on Unstructured Streams of DataMatthew Russo, Tatsunori Hashimoto, Daniel Kang, Yi Sun et al.VLDB 2023 · 10 citations
- Co-movement Pattern Mining from VideosDongxiang Zhang, Teng Ma, Junnan Hu, Yijun Bei et al.VLDB 2024 · 8 citations
- SketchQL: Video Moment Querying with a Visual Query InterfaceRenzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu et al.SIGMOD 2025 · 6 citations
Builds on1
Related papers
- Ranked Window Query Retrieval over Video RepositoriesYueting Chen, Xiaohui Yu, Nick KoudasICDE 2022 · 9 citations
- Context-Aware Relative Object Queries to Unify Video Instance and Panoptic SegmentationAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingCVPR 2023
- Track Merging for Effective Video Query ProcessingDaren Chao, Yueting Chen, Nick Koudas, Xiaohui YuICDE 2023 · 5 citations
- Learning Spatial-Semantic Features for Robust Video Object SegmentationXin Li, Deshui Miao, Zhenyu He, Yaowei Wang et al.ICLR 2025
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See et al.AAAI 2020 · 18 citations
