GEqO: ML-Accelerated Semantic Equivalence Detection
Brandon Haynes, Rana Alotaibi, Anna Pavlenko, Jyoti Leeka, Alekh Jindal, Yuanyuan Tian
摘要
Large scale analytics engines have become a core dependency for modern data-driven enterprises to derive business insights and drive actions. These engines support a large number of analytic jobs processing huge volumes of data on a daily basis, and workloads are often inundated with overlapping computations across multiple jobs. Reusing common computation is crucial for efficient cluster resource utilization and reducing job execution time. Detecting common computation is the first and key step for reducing this computational redundancy. However, detecting equivalence on large-scale analytics engines requires efficient and scalable solutions that are fully automated. In addition, to maximize computation reuse, equivalence needs to be detected at the semantic level instead of just the syntactic level (i.e., the ability to detect semantic equivalence of seemingly different-looking queries). Unfortunately, existing solutions fall short of satisfying these requirements. In this paper, we take a major step towards filling this gap by proposing GEqO, a portable and lightweight machine-learning-based framework for efficiently identifying semantically equivalent computations at scale. GEqO introduces two machine-learning-based filters that quickly prune out nonequivalent subexpressions and employs a semi-supervised learning feedback loop to iteratively improve its model with an intelligent sampling mechanism. Further, with its novel database-agnostic featurization method, GEqO can transfer the learning from one workload and database to another. Our extensive empirical evaluation shows that, on TPC-DS-like queries, GEqO yields significant performance gains-up to 200x faster than automated verifiers-and finds up to 2x more equivalences than optimizer and signature-based equivalence detection approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mayura: Exploiting Similarities in Motifs for Temporal Co-MiningSanjay Sri Vallabh Singapuram, Ronald G. Dreslinski, Nishil TalatiVLDB 2025
- TATA: An Efficient Framework for Task Transfer in Query Plan RepresentationYue Zhao, Songsong Mo, Gao CongVLDB 2026
它引用的顶会 Paper6
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul 等SIGMOD 2021 · 被引用 242 次
- Automatic View Generation with Deep Learning and Reinforcement LearningHaitao Yuan, Guoliang Li, Ling Feng, Ji Sun 等ICDE 2020 · 被引用 66 次
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- Automatic Detection of Performance Bugs in Database Systems using Equivalent QueriesXinyu Liu, Qi Zhou, Joy Arulraj, Alessandro OrsoICSE 2022 · 被引用 42 次
- WeTune: Automatic Discovery and Verification of Query Rewrite RulesZhaoguo Wang, Zhou Zhou, Yicun Yang, Haoran Ding 等SIGMOD 2022 · 被引用 35 次
相关 Paper
- Auto-Differentiation of Relational Computations for Very Large Scale Machine LearningYuxin Tang, Zhimin Ding, Dimitrije Jankov, Binhang Yuan 等ICML 2023 · 被引用 7 次
- EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action PruningJiawei Liu, Qisi Chen, Jianshu Zhang, Quan Liu 等ACL 2026 · 被引用 1 次
- Practical Parameterized Query Optimization via Efficient Plan Reuse and List-wise RankingHai Lan, Yang Yu, Zhifeng Bao, Zi Huang 等SIGMOD 2026
- HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics QueriesRana Alotaibi, Bogdan Cautis, Alin Deutsch, Ioana ManolescuSIGMOD 2021 · 被引用 4 次
- SPES: A Symbolic Approach to Proving Query Equivalence Under Bag SemanticsQi Zhou, Joy Arulraj, Shamkant B. Navathe, William Harris 等ICDE 2022 · 被引用 19 次
