OTIF: Efficient Tracker Pre-processing over Large Video Datasets
Favyen Bastani, Samuel Madden
摘要
Performing analytics tasks over large-scale video datasets is increasingly common in a wide range of applications, from traffic planning to sports analytics. These tasks generally involve object detection and tracking operations that require pre-processing the video through expensive machine learning models. To address this cost, several video query optimizers have recently been proposed. Broadly, these methods trade large reductions in pre-processing cost for increases in query execution cost: during query execution, they apply query-specific machine learning operations over portions of the video dataset. Although video query optimizers reduce the overall cost of executing a single query over large video datasets compared to naive object tracking methods, executing several queries over the same video remains cost-prohibitive; moreover, the high per-query latency makes these systems unsuitable for exploratory analytics where fast response times are crucial.
In this paper, we present OTIF, a video pre-processor that efficiently extracts all object tracks from large-scale video datasets. By integrating several optimizations under a joint parameter tuning framework, OTIF is able to extract all object tracks from video as fast as existing video query optimizers can execute just one single query. In contrast to the outputs of video query optimizers, OTIF's outputs are general-purpose object tracks that can be used to execute many queries with sub-second latencies. We compare OTIF against three recent video query optimizers, as well as several general-purpose object detection and tracking techniques, and find that, across multiple datasets, OTIF provides a 6x to 25x average reduction in the overall cost to execute five queries over the same video.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Extract-Transform-Load for Video StreamsFerdinand Kossmann, Ziniu Wu, Eugenie Lai, Nesime Tatbul 等VLDB 2023 · 被引用 21 次
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes 等VLDB 2023 · 被引用 18 次
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 被引用 11 次
- Predictive and Near-Optimal Sampling for View Materialization in Video DatabasesYanchao Xu, Dongxiang Zhang, Shuhao Zhang, Sai Wu 等SIGMOD 2024 · 被引用 5 次
- Optimizing Video Queries with Declarative CluesDaren Chao, Yueting Chen, Nick Koudas, Xiaohui YuVLDB 2024 · 被引用 5 次
它引用的顶会 Paper5
- FAMNet: Joint Learning of Feature, Affinity and Multi-Dimensional Assignment for Online Multiple Object TrackingPeng Chu, Haibin LingICCV 2019 · 被引用 229 次
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 被引用 143 次
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan 等SIGMOD 2020 · 被引用 68 次
- Video Monitoring QueriesNick Koudas, Raymond Li, Ioannis XarchakosICDE 2020 · 被引用 33 次
- TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured DataDaniel Kang, John Guibas, Peter D. Bailis, Tatsunori Hashimoto 等SIGMOD 2022 · 被引用 23 次
相关 Paper
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu 等ICDE 2025 · 被引用 2 次
- QaVA: Query-Aware Video Analysis Framework Based on Data Access PatternTianxiong Zhong, Zhiwei Zhang, Yihang Fu, Guo Lu 等ICDE 2025 · 被引用 1 次
- Track Merging for Effective Video Query ProcessingDaren Chao, Yueting Chen, Nick Koudas, Xiaohui YuICDE 2023 · 被引用 5 次
- Context-Aware Relative Object Queries to Unify Video Instance and Panoptic SegmentationAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingCVPR 2023
- Craw: A Unified and Efficient Querying Framework for Large-Scale Video DatasetsZiqi Zhou, Hanjian Jiang, Zihao Zeng, Xupuzhe Shao 等VLDB 2026
