TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data
Daniel Kang, John Guibas, Peter D. Bailis, Tatsunori Hashimoto, Matei Zaharia
摘要
Unstructured data (e.g., video or text) is now commonly queried by using computationally expensive deep neural networks or human labelers to produce structured information, e.g., object types and positions in video. To accelerate queries, many recent systems (e.g., BlazeIt, NoScope, Tahoma, SUPG, etc.) train a query-specific proxy model to approximate a large target labelers (i.e., these expensive neural networks or human labelers). These models return proxy scores that are then used in query processing algorithms. Unfortunately, proxy models usually have to be trained per query and require large amounts of annotations from the target labelers.
In this work, we develop an index (trainable semantic index, TASTI) that simultaneously removes the need for per-query proxies and is more efficient to construct than prior indexes. TASTI accomplishes this by leveraging semantic similarity across records in a given dataset. Specifically, it produces embeddings for each record such that records with close embeddings have similar target labeler outputs. TASTI then generates high-quality proxy scores via embeddings without needing to train a per-query proxy. These scores can be used in existing proxy-based query processing algorithms (e.g., for aggregation, selection, etc.). We theoretically analyze TASTI and show that a low embedding training error guarantees downstream query accuracy for a natural class of queries. We evaluate TASTI on five video, text, and speech datasets, and three query types. We show that TASTI's indexes can be 10× less expensive to construct than generating annotations for current proxy-based methods, and accelerate queries by up to 24×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Privid: Practical, Privacy-Preserving Video Analytics QueriesFrank Cangialosi, Neil Agarwal, Venkat Arun, Junchen Jiang 等NSDI 2022 · 被引用 36 次
- Extract-Transform-Load for Video StreamsFerdinand Kossmann, Ziniu Wu, Eugenie Lai, Nesime Tatbul 等VLDB 2023 · 被引用 21 次
- OTIF: Efficient Tracker Pre-processing over Large Video DatasetsFavyen Bastani, Samuel MaddenSIGMOD 2022 · 被引用 20 次
- EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User InteractionsEnhao Zhang, Maureen Daum, Dong He, Brandon Haynes 等VLDB 2023 · 被引用 18 次
- ThalamusDB: Approximate Query Processing on Multi-Modal DataSaehan Jo, Immanuel TrummerSIGMOD 2024 · 被引用 11 次
它引用的顶会 Paper6
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- MIRIS: Fast Object Track Queries in VideoFavyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan 等SIGMOD 2020 · 被引用 68 次
- Jointly Optimizing Preprocessing and Inference for DNN-based Visual AnalyticsDaniel Kang, Ankit Mathur, Teja Veeramacheneni, Peter Bailis 等VLDB 2021 · 被引用 50 次
- Approximate Selection with Guarantees using ProxiesDaniel Kang, Edward Gan, Peter Bailis, Tatsunori Hashimoto 等VLDB 2020 · 被引用 46 次
- Accelerating Approximate Aggregation Queries with Expensive PredicatesDaniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto 等VLDB 2021 · 被引用 34 次
相关 Paper
- HAIDES: Adaptive Approximation of Inference Queries over Unstructured DataChristos C. Papadopoulos, Alkis Simitsis, Torben Bach PedersenICDE 2025
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis]Yeounoh Chung, Rushabh Desai, Jian He, Yu Xiao 等SIGMOD 2026 · 被引用 8 次
- QaVA: Query-Aware Video Analysis Framework Based on Data Access PatternTianxiong Zhong, Zhiwei Zhang, Yihang Fu, Guo Lu 等ICDE 2025 · 被引用 1 次
- Accelerating Aggregation Queries on Unstructured Streams of DataMatthew Russo, Tatsunori Hashimoto, Daniel Kang, Yi Sun 等VLDB 2023 · 被引用 10 次
- LOVO: Efficient Complex Object Query in Large-Scale Video DatasetsYuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu 等ICDE 2025 · 被引用 2 次
