Optimizing Video Selection LIMIT Queries With Commonsense Knowledge
Wenjia He, Ibrahim Sabek, Yuze Lou, Michael J. Cafarella
摘要
Video is becoming a major part of contemporary data collection. It is increasingly important to process video selection queries --- selecting videos that contain target objects. Advances in neural networks allow us to detect the objects in an image, and thereby offer query systems to examine the content of the video. Unfortunately, neural network-based approaches have long inference times. Processing this type of query through a standard scan would be time-consuming and would involve applying complex detectors to numerous irrelevant videos. It is tempting to try to improve query times by computing an index in advance. But unfortunately, many frames will never be beneficial for any query. Time spent processing them, whether at index time or at query time, is simply wasted computation.
We propose a novel index mechanism to optimize video selection queries with commonsense knowledge. Commonsense knowledge consists of fundamental information about the world, such as the fact that a tennis racket is a tool designed for hitting a tennis ball. To save computation, an inexpensive but lossy index can be intentionally created, but this may result in missed target objects and suboptimal query time performance. Our mechanism addresses this issue by constructing probabilistic models from commonsense knowledge to patch the lossy index and then prioritizing predicate-related videos at query time. This method can achieve significant performance improvements comparable to those of a full index while keeping the construction costs of a lossy index. We describe our prototype system, Paine, plus experiments on two video corpora. We show our best optimization method can process up to 97.79% fewer videos compared to baselines. Even the model constructed without any video content can yield a 75.39% improvement over baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan 等SIGMOD 2020 · 被引用 68 次
- Jointly Optimizing Preprocessing and Inference for DNN-based Visual AnalyticsDaniel Kang, Ankit Mathur, Teja Veeramacheneni, Peter Bailis 等VLDB 2021 · 被引用 50 次
相关 Paper
- ExSample: Efficient Searches on Video Repositories through Adaptive SamplingOscar R. Moll, Favyen Bastani, Sam Madden, Mike Stonebraker 等ICDE 2022 · 被引用 16 次
- QaVA: Query-Aware Video Analysis Framework Based on Data Access PatternTianxiong Zhong, Zhiwei Zhang, Yihang Fu, Guo Lu 等ICDE 2025 · 被引用 1 次
- Joint Commonsense and Relation Reasoning for Image and Video CaptioningJingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi 等AAAI 2020 · 被引用 52 次
- Commonsense for Zero-Shot Natural Language Video LocalizationMeghana Holla, Ismini LourentzouAAAI 2024 · 被引用 6 次
- FiGO: Fine-Grained Query Optimization in Video AnalyticsJiashen Cao, Karan Sarkar, Ramyad Hadidi, Joy Arulraj 等SIGMOD 2022 · 被引用 37 次
