Optimizing Machine Learning Inference Queries with Correlative Proxy Models
Zhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu, Chen Li, X. Sean Wang
Abstract
We consider accelerating machine learning (ML) inference queries on unstructured datasets. Expensive operators such as feature extractors and classifiers are deployed as user-defined functions (UDFs), which are not penetrable with classic query optimization techniques such as predicate push-down. Recent optimization schemes (e.g., Probabilistic Predicates or PP) assume independence among the query predicates, build a proxy model for each predicate offline, and rewrite a new query by injecting these cheap proxy models in the front of the expensive ML UDFs. In such a manner, unlikely inputs that do not satisfy query predicates are filtered early to bypass the ML UDFs. We show that enforcing the independence assumption in this context may result in sub-optimal plans. In this paper, we propose CORE, a query optimizer that better exploits the predicate correlations and accelerates ML inference queries. Our solution builds the proxy models online for a new query and leverages a branch-and-bound search process to reduce the building costs. Results on three real-world text, image and video datasets show that CORE improves the query throughput by up to 63% compared to PP and up to 80% compared to running the queries as it is.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df14f7b6-b135-4ea0-b27f-6647e3bee975Cited by top-tier papers12
- Optimizing Video Analytics with Declarative Model RelationshipsFrancisco Romero, Johann Hauswald, Aditi Partap, Daniel Kang et al.VLDB 2023 · 37 citations
- Serving and Optimizing Machine Learning Workflows on Heterogeneous InfrastructuresYongji Wu, Matthew Lentz, Danyang Zhuo, Yao LuVLDB 2023 · 31 citations
- Extract-Transform-Load for Video StreamsFerdinand Kossmann, Ziniu Wu, Eugenie Lai, Nesime Tatbul et al.VLDB 2023 · 21 citations
- SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained EnvironmentsQiuru Lin, Sai Wu, Junbo Zhao, Jian Dai et al.VLDB 2024 · 17 citations
- On Efficient Approximate Queries over Machine Learning ModelsDujian Ding, Sihem Amer-Yahia, Laks V. S. LakshmananVLDB 2023 · 11 citations
Builds on3
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 103 citations
- Approximate Selection with Guarantees using ProxiesDaniel Kang, Edward Gan, Peter Bailis, Tatsunori Hashimoto et al.VLDB 2020 · 46 citations
- Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource ConstraintsShaofeng Cai, Gang Chen, Beng Chin Ooi, Jinyang GaoVLDB 2020 · 21 citations
Related papers
- PLAQUE: Automated Predicate Learning at Query TimeYiming Lin, Sharad MehrotraSIGMOD 2024 · 2 citations
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 15 citations
- A Method for Optimizing Opaque Filter QueriesWenjia He, Michael R. Anderson, Maxwell Strome, Michael J. CafarellaSIGMOD 2020 · 15 citations
- Mitigating the Impedance Mismatch between Prediction Query Execution and Database EngineChenyang Zhang, Junxiong Peng, Chen Xu, Quanqing Xu et al.SIGMOD 2025 · 7 citations
- Aero: Adaptive Query Processing of ML QueriesGaurav Tarlok Kakkar, Jiashen Cao, Aubhro Sengupta, Joy Arulraj et al.SIGMOD 2025 · 2 citations
