LOVO: Efficient Complex Object Query in Large-Scale Video Datasets
Yuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu, Xinyu Lu, Jiong Lou, Chentao Wu, Wei Zhao, Jie Li
Abstract
The widespread deployment of cameras has led to an exponential increase in video data, creating vast opportunities for applications such as traffic management and crime surveillance. However, querying specific objects from large-scale video datasets presents challenges, including (1) processing massive and continuously growing data volumes, (2) supporting complex query requirements, and (3) ensuring low-latency execution. Existing video analysis methods struggle with either limited adaptability to unseen object classes or suffer from high query latency.
In this paper, we present LOVO, a novel system designed to efficiently handle compLex Object queries in large-scale VideO datasets. Agnostic to user queries, LOVO performs one-time feature extraction using pre-trained visual encoders, generating compact visual embeddings for key frames to build an efficient index. These visual embeddings, along with associated bounding boxes, are organized in an inverted multi-index structure within a vector database, which supports queries for any objects. During the query phase, LOVO transforms object queries to query embeddings and conducts fast approximate nearest-neighbor searches on the visual embeddings. Finally, a cross-modal rerank is performed to refine the results by fusing visual features with detailed textual features. Evaluation on real-world video datasets demonstrates that LOVO outperforms existing methods in handling complex queries, with near-optimal query accuracy and up to 85x lower search latency, while significantly reducing index construction costs. This system redefines the state-of-theart object query approaches in video analysis, setting a new benchmark for complex object queries with a novel, scalable, and efficient approach that excels in dynamic environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1e349f0-3446-4abb-a976-8be9a3a3c720Cited by top-tier papers1
Ask how each one uses itBuilds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen et al.CVPR 2022 · 150 citations
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang et al.NeurIPS 2024 · 147 citations
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 103 citations
Related papers
- OTIF: Efficient Tracker Pre-processing over Large Video DatasetsFavyen Bastani, Samuel MaddenSIGMOD 2022 · 20 citations
- Lava: Language Driven Scalable and Versatile Traffic Video AnalyticsYanrui Yu, Tianfei Zhou, Jiaxin Sun, Lianpeng Qiao et al.ACM MM 2025
- ARC: Approximate Relevant Clip Query in Large-Scale Video RepositoriesYue Chen, Yinan Jing, Ziqiang Yu, Xiaohui Yu et al.SIGIR 2025 · 1 citation
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang et al.SIGIR 2020 · 131 citations
- QaVA: Query-Aware Video Analysis Framework Based on Data Access PatternTianxiong Zhong, Zhiwei Zhang, Yihang Fu, Guo Lu et al.ICDE 2025 · 1 citation
