OTIF: Efficient Tracker Pre-processing over Large Video Datasets
Favyen Bastani, Samuel Madden
Abstract
Performing analytics tasks over large-scale video datasets is increasingly common in a wide range of applications, from traffic planning to sports analytics. These tasks generally involve object detection and tracking operations that require pre-processing the video through expensive machine learning models. To address this cost, several video query optimizers have recently been proposed. Broadly, these methods trade large reductions in pre-processing cost for increases in query execution cost: during query execution, they apply query-specific machine learning operations over portions of the video dataset. Although video query optimizers reduce the overall cost of executing a single query over large video datasets compared to naive object tracking methods, executing several queries over the same video remains cost-prohibitive; moreover, the high per-query latency makes these systems unsuitable for exploratory analytics where fast response times are crucial.
In this paper, we present OTIF, a video pre-processor that efficiently extracts all object tracks from large-scale video datasets. By integrating several optimizations under a joint parameter tuning framework, OTIF is able to extract all object tracks from video as fast as existing video query optimizers can execute just one single query. In contrast to the outputs of video query optimizers, OTIF's outputs are general-purpose object tracks that can be used to execute many queries with sub-second latencies. We compare OTIF against three recent video query optimizers, as well as several general-purpose object detection and tracking techniques, and find that, across multiple datasets, OTIF provides a 6x to 25x average reduction in the overall cost to execute five queries over the same video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dc9d376-2b87-4105-93c4-90b9ec6d9639Cited by top-tier papers13
- Extract-Transform-Load for Video StreamsFerdinand Kossmann, Ziniu Wu, Eugenie Lai, Nesime Tatbul et al.VLDB 2023 · 21 citations
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes et al.VLDB 2023 · 18 citations
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 11 citations
- Predictive and Near-Optimal Sampling for View Materialization in Video DatabasesYanchao Xu, Dongxiang Zhang, Shuhao Zhang, Sai Wu et al.SIGMOD 2024 · 5 citations
- Optimizing Video Queries with Declarative CluesDaren Chao, Yueting Chen, Nick Koudas, Xiaohui YuVLDB 2024 · 5 citations
Builds on5
- FAMNet: Joint Learning of Feature, Affinity and Multi-Dimensional Assignment for Online Multiple Object TrackingPeng Chu, Haibin LingICCV 2019 · 229 citations
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 143 citations
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan et al.SIGMOD 2020 · 68 citations
- Video Monitoring QueriesNick Koudas, Raymond Li, Ioannis XarchakosICDE 2020 · 33 citations
- TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured DataDaniel Kang, John Guibas, Peter D. Bailis, Tatsunori Hashimoto et al.SIGMOD 2022 · 23 citations
Related papers
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu et al.ICDE 2025 · 2 citations
- QaVA: Query-Aware Video Analysis Framework Based on Data Access PatternTianxiong Zhong, Zhiwei Zhang, Yihang Fu, Guo Lu et al.ICDE 2025 · 1 citation
- Track Merging for Effective Video Query ProcessingDaren Chao, Yueting Chen, Nick Koudas, Xiaohui YuICDE 2023 · 5 citations
- Context-Aware Relative Object Queries to Unify Video Instance and Panoptic SegmentationAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingCVPR 2023
- Craw: A Unified and Efficient Querying Framework for Large-Scale Video DatasetsZiqi Zhou, Hanjian Jiang, Zihao Zeng, Xupuzhe Shao et al.VLDB 2026
