SIA: Optimizing Queries using Learned Predicates
Qi Zhou, Joy Arulraj, Shamkant B. Navathe, William Harris, Jinpeng Wu
Abstract
Predicate-centric rules for rewriting queries is a key technique in optimizing queries. These include pushing down the predicate below the join and aggregation operators, or optimizing the order of evaluating predicates. However, many of these rules are only applicable when the predicate uses a certain set of columns. For example, to move the predicate below the join operator, the predicate must only use columns from one of the joined tables. By generating a predicate that satisfies these column constraints and preserves the semantics of the original query, the optimizer may leverage additional predicate-centric rules that were not applicable before. Researchers have proposed syntax-driven rewrite rules and machine learning algorithms for inferring such predicates. However, these techniques suffer from two limitations. First, they do not let the optimizer constrain the set of columns that may be used in the learned predicate. Second, machine learning algorithms do not guarantee that the learned predicate preserves semantics. In this paper, we present SIA, a system for learning predicates while being guided by counter-examples and a verification technique, that addresses these limitations. The key idea is to leverage satisfiability modulo theories to generate counter-examples and use them to iteratively learn a valid, optimal predicate. We formalize this problem by proving the key properties of synthesized predicates. We implement our approach in SIA and evaluate its efficacy and efficiency. We demonstrate that it synthesizes a larger set of valid predicates compared to prior approaches. On a collection of 200 queries derived from the TPC-H benchmark, SIA successfully rewrites 114 queries with learned predicates. 66 of these rewritten queries exhibit more than 2X speed up.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 84c4180f-2edd-4113-8a8c-55a47bbc274aCited by top-tier papers8
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 15 citations
- SlabCity: Whole-Query Optimization using Program SynthesisRui Dong, Jie Liu, Yuxuan Zhu, Cong Yan et al.VLDB 2023 · 11 citations
- AgenticScholar: Agentic Data Management with Pipeline Orchestration for Scholarly CorporaHai Lan, Tingting Wang, Zhifeng Bao, Guoliang Li et al.SIGMOD 2026 · 4 citations
- Verifying Data Constraint Equivalence in FinTech SystemsChengpeng Wang, Gang Fan, Peisen Yao, Fuxiong Pan et al.ICSE 2023 · 4 citations
- TRAP: Tailored Robustness Assessment for Index Advisors via Adversarial PerturbationWei Zhou, Chen Lin, Xuanhe Zhou, Guoliang Li et al.ICDE 2024 · 3 citations
Related papers
- PLAQUE: Automated Predicate Learning at Query TimeYiming Lin, Sharad MehrotraSIGMOD 2024 · 2 citations
- Relational Query Synthesis ⋈ Decision Tree LearningAaditya Naik, Aalok Thakkar, Adam Stein, Rajeev Alur et al.VLDB 2024 · 2 citations
- WeTune: Automatic Discovery and Verification of Query Rewrite RulesZhaoguo Wang, Zhou Zhou, Yicun Yang, Haoran Ding et al.SIGMOD 2022 · 35 citations
- Mobius: Synthesizing Relational Queries with Recursive and Invented PredicatesAalok Thakkar, Nathaniel Sands, George Petrou, Rajeev Alur et al.OOPSLA 2023 · 5 citations
- Provenance-guided synthesis of Datalog programsMukund Raghothaman, Jonathan Mendelson, David Zhao, Mayur Naik et al.POPL 2020 · 49 citations
