COMPARE: Accelerating Groupwise Comparison in Relational Databases for Data Analytics
Tarique Siddiqui, Surajit Chaudhuri, Vivek R. Narasayya
摘要
Data analysis often involves comparing subsets of data across many dimensions for finding unusual trends and patterns. While the comparison between subsets of data can be expressed using SQL, they tend to be complex to write, and suffer from poor performance over large and high-dimensional datasets. In this paper, we propose a new logical operator COMPARE for relational databases that concisely captures the enumeration and comparison between subsets of data and greatly simplifies the expressing of a large class of comparative queries. We extend the database engine with optimization techniques that exploit the semantics of COMPARE to significantly improve the performance of such queries. We have implemented these extensions inside Microsoft SQL Server, a commercial DBMS engine. Our extensive evaluation on synthetic and real-world datasets shows that COMPARE results in a significant speedup over existing approaches, including physical plans generated by today's database systems, user-defined functions (UDFs), as well as middleware solutions that compare subsets outside the databases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Incorporating Super-Operators in Big-Data Query OptimizersJyoti Leeka, Kaushik RajanVLDB 2020 · 被引用 17 次
- Procedural Extensions of SQL: Understanding their usage in the wildSurabhi Gupta, Karthik RamachandraVLDB 2021 · 被引用 34 次
- Database Technology for the Masses: Sub-Operators as First-Class EntitiesMaximilian Bandle, Jana GicevaVLDB 2021 · 被引用 19 次
- "What makes my queries slow?": Subgroup Discovery for SQL Workload AnalysisYoucef Remil, Anes Bendimerad, Romain Mathonat, Philippe Chaleat 等ASE 2021 · 被引用 12 次
- Computing the Difference of Conjunctive Queries EfficientlyXiao Hu, Qichen WangSIGMOD 2023 · 被引用 9 次
