Accelerating Aggregation Queries on Unstructured Streams of Data
Matthew Russo, Tatsunori Hashimoto, Daniel Kang, Yi Sun, Matei Zaharia
摘要
Analysts and scientists are interested in querying streams of video, audio, and text to extract quantitative insights. For example, an urban planner may wish to measure congestion by querying the live feed from a traffic camera. Prior work has used deep neural networks (DNNs) to answer such queries in the batch setting. However, much of this work is not suited for the streaming setting because it requires access to the entire dataset before a query can be submitted or is specific to video. Thus, to the best of our knowledge, no prior work addresses the problem of efficiently answering queries over multiple modalities of streams. In this work we propose InQuest, a system for accelerating aggregation queries on unstructured streams of data with statistical guarantees on query accuracy. InQuest leverages inexpensive approximation models ("proxies") and sampling techniques to limit the execution of an expensive high-precision model (an "oracle") to a subset of the stream. It then uses the oracle predictions to compute an approximate query answer in real-time. We theoretically analyzed InQuest and show that the expected error of its query estimates converges on stationary streams at a rate inversely proportional to the oracle budget. We evaluated our algorithm on six real-world video and text datasets and show that InQuest achieves the same root mean squared error (RMSE) as two streaming baselines with up to 5.0x fewer oracle invocations. We further show that InQuest can achieve up to 1.9x lower RMSE at a fixed number of oracle invocations than a state-of-the-art batch setting algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano 等VLDB 2026 · 被引用 9 次
- PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error GuaranteesYuxuan Zhu, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang 等SIGMOD 2025 · 被引用 3 次
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim 等VLDB 2025 · 被引用 2 次
它引用的顶会 Paper10
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan 等SIGMOD 2020 · 被引用 68 次
- Jointly Optimizing Preprocessing and Inference for DNN-based Visual AnalyticsDaniel Kang, Ankit Mathur, Teja Veeramacheneni, Peter Bailis 等VLDB 2021 · 被引用 50 次
- Approximate Selection with Guarantees using ProxiesDaniel Kang, Edward Gan, Peter Bailis, Tatsunori Hashimoto 等VLDB 2020 · 被引用 46 次
- Accelerating Approximate Aggregation Queries with Expensive PredicatesDaniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto 等VLDB 2021 · 被引用 34 次
相关 Paper
- HAIDES: Adaptive Approximation of Inference Queries over Unstructured DataChristos C. Papadopoulos, Alkis Simitsis, Torben Bach PedersenICDE 2025
- Video Monitoring QueriesNick Koudas, Raymond Li, Ioannis XarchakosICDE 2020 · 被引用 33 次
- On Efficient Approximate Aggregate Nearest Neighbor Queries over Learned RepresentationsCarrie Wang, Sihem Amer-Yahia, Laks V. S. Lakshmanan, Reynold ChengSIGMOD 2026
- SEIDEN: Revisiting Query Processing in Video Database SystemsJaeho Bang, Gaurav Tarlok Kakkar, Pramod Chunduri, Subrata Mitra 等VLDB 2023 · 被引用 24 次
- Top-K Deep Video Analytics: A Probabilistic ApproachZiliang Lai, Chenxia Han, Chris Liu, Pengfei Zhang 等SIGMOD 2021 · 被引用 7 次
