On Efficient Approximate Queries over Machine Learning Models
Dujian Ding, Sihem Amer-Yahia, Laks V. S. Lakshmanan
摘要
The question of answering queries over ML predictions has been gaining attention in the database community. This question is challenging because finding high quality answers by invoking an oracle such as a human expert or an expensive deep neural network model on every single item in the DB and then applying the query, can be prohibitive. We develop a novel unified framework for approximate query answering by leveraging a proxy to minimize the oracle usage of finding high quality answers for both Precision-Target (PT) and Recall-Target (RT) queries. Our framework uses a judicious combination of invoking the expensive oracle on data samples and applying the cheap proxy on the DB objects. It relies on two assumptions. Under the P roxy Q uality assumption, we develop two algorithms: PQA that efficiently finds high quality answers with high probability and no oracle calls, and PQE, a heuristic extension that achieves empirically good performance with a small number of oracle calls. Alternatively, under the C ore S et C losure assumption, we develop two algorithms: CSC that efficiently returns high quality answers with high probability and minimal oracle usage, and CSE, which extends it to more general settings. Our extensive experiments on five real-world datasets on both query types, PT and RT, demonstrate that our algorithms outperform the state-of-the-art and achieve high result quality with provable statistical guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim 等ICLR 2024 · 被引用 282 次
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 被引用 11 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
- Cut Costs, Not Accuracy: LLM-Powered Data Processing with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranSIGMOD 2026 · 被引用 1 次
- OCCAM: Towards Cost-Efficient and Accuracy-Aware Classification InferenceDujian Ding, Bicheng Xu, Laks V. S. LakshmananICLR 2025
它引用的顶会 Paper5
- Approximate Selection with Guarantees using ProxiesDaniel Kang, Edward Gan, Peter Bailis, Tatsunori Hashimoto 等VLDB 2020 · 被引用 46 次
- Optimizing Machine Learning Inference Queries with Correlative Proxy ModelsZhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu 等VLDB 2022 · 被引用 33 次
- Top-K Deep Video Analytics: A Probabilistic ApproachZiliang Lai, Chenxia Han, Chris Liu, Pengfei Zhang 等SIGMOD 2021 · 被引用 7 次
- Efficiently Answering Durability Prediction QueriesJunyang Gao, Yifan Xu, Pankaj K. Agarwal, Jun YangSIGMOD 2021 · 被引用 3 次
- A Generalized Approach for Reducing Expensive Distance Calls for A Broad Class of Proximity ProblemsJees Augustine, Suraj Shetiya, Mohammadreza Esfandiari, Senjuti Basu Roy 等SIGMOD 2021 · 被引用 1 次
相关 Paper
- Robust Plan Evaluation based on Approximate Probabilistic Machine LearningAmin Kamali, Verena Kantere, Calisto Zuzarte, Vincent CorvinelliVLDB 2025 · 被引用 1 次
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 被引用 54 次
- Accelerating Approximate Aggregation Queries with Expensive PredicatesDaniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto 等VLDB 2021 · 被引用 34 次
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis]Yeounoh Chung, Rushabh Desai, Jian He, Yu Xiao 等SIGMOD 2026 · 被引用 8 次
- Consistent Range Approximation for Fair Predictive ModelingJiongli Zhu, Sainyam Galhotra, Nazanin Sabri, Babak SalimiVLDB 2023 · 被引用 14 次
