Parallel Query Processing: To Separate Communication from Computation
Hao Zhang, Jeffrey Xu Yu, Yikai Zhang, Kangfei Zhao
摘要
In this paper, we study parallel query processing with a focus on reducing the communication cost, which is the dominating factor in parallel query processing. The communication cost becomes large if the intermediate results between operators are large in intra-operator parallelism. In the existing approaches, it optimizes an SQL query by arranging relational algebra operators to reduce the total cost, where, for each operator, it involves (i) distribution of data partitioned to computing nodes by communication, and (ii)computation on computing nodes locally. The communication and computation are dealt with inside an operator and are not separable. In other words, it is difficult to avoid large intermediate results and hence reduce the communication cost. To reduce communication cost, we separate communication from computation using several new operators proposed in this paper. One is a pair operator () to pair the partitions of a relation R with the partitions of a relation S, where a partition is specified by a hash function. With the pair operator defined, we can explicitly deal with communication to deliver pairs of partitions to computing nodes. Together with , we can also explicitly treat the local computation on a computing node as op for any RA (relational algebra) operator op. We give a merge operator (U), to collect all partial results from computing nodes as they are. In short, with , op, and U, we are able to explicitly specify communication and computation for RA operators. Furthermore, we propose new techniques, namely, partitioning push-down and computation push-up to separate communication from computation for RA expressions. We prove that we can push-down/up for a wide range of relational expressions. We have developed a distributed system named Secco (Separate Communication from Computation) by revamping SparkSQL on Spark, and confirmed the efficiency of our approach in our performance studies using real datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Distributed Evaluation of Graph Queries Using Recursive Relational AlgebraSarah Chlyah, Pierre Genevès, Nabil LayaïdaICDE 2025
- New Query Optimization Techniques in the Spark Engine of Azure SynapseAbhishek Modi, Kaushik Rajan, Srinivas Thimmaiah, Prakhar Jain 等VLDB 2022 · 被引用 13 次
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
- Thrifty Query Execution via IncrementabilityDixin Tang, Zechao Shang, Aaron J. Elmore, Sanjay Krishnan 等SIGMOD 2020 · 被引用 9 次
- LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise ComparisonJunhao Ye, Jiahui Li, Lu Chen, Yuren Mao 等VLDB 2025 · 被引用 2 次
