HAIDES: Adaptive Approximation of Inference Queries over Unstructured Data
Christos C. Papadopoulos, Alkis Simitsis, Torben Bach Pedersen
Abstract
Modern analytics rely on insights derived from the execution of inference queries over vast amounts of unstructured data such as text, images, and video. Oftentimes, these queries evaluate predicates based on an expensive “oracle“ model in the likes of a deep neural network or human input that dominates the total query cost. Prior work has focused on training computationally cheap proxy models at query time that produce an approximate result. Alternatively, index-based methods apply the original oracle over a representative set of data points and generate the approximate result through an inference propagation process. Current state-of-the-art (SOTA) index-based methods require a memory -expensive index construction process which offsets their oracle cost-effectiveness and can make their usage prohibitive. In this work, we present HAIDES, an index-based, domain-agnostic framework for approximating inference on unstructured data. HAIDES consists of two main components: a coarse-to-fine framework that can be efficiently constructed using minimal memory, and a novel index adaptation component that makes use of oracle invocations during query execution in order to adaptively produce representative sets that yield high-quality approximate results. Our experimental results across three challenging domains-video, images, text-show that HAIDES (a) constructs indexes that produce performant representative sets with up to 2 orders of magnitude less memory than the SOTA baseline, while (b) requires up to 2x less oracle calls to produce the same result quality, and (c) achieves up to 10 percentage points better result quality when using the same oracle calls.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7d0ec942-4bde-4bee-96c3-246be54cef5fRelated papers
- Accelerating Aggregation Queries on Unstructured Streams of DataMatthew Russo, Tatsunori Hashimoto, Daniel Kang, Yi Sun et al.VLDB 2023 · 10 citations
- TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured DataDaniel Kang, John Guibas, Peter D. Bailis, Tatsunori Hashimoto et al.SIGMOD 2022 · 23 citations
- Approximate Selection with Guarantees using ProxiesDaniel Kang, Edward Gan, Peter Bailis, Tatsunori Hashimoto et al.VLDB 2020 · 46 citations
- SEIDEN: Revisiting Query Processing in Video Database SystemsJaeho Bang, Gaurav Tarlok Kakkar, Pramod Chunduri, Subrata Mitra et al.VLDB 2023 · 24 citations
- Optimizing Machine Learning Inference Queries with Correlative Proxy ModelsZhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu et al.VLDB 2022 · 33 citations
