Parallel Query Processing: To Separate Communication from Computation
Hao Zhang, Jeffrey Xu Yu, Yikai Zhang, Kangfei Zhao
Abstract
In this paper, we study parallel query processing with a focus on reducing the communication cost, which is the dominating factor in parallel query processing. The communication cost becomes large if the intermediate results between operators are large in intra-operator parallelism. In the existing approaches, it optimizes an SQL query by arranging relational algebra operators to reduce the total cost, where, for each operator, it involves (i) distribution of data partitioned to computing nodes by communication, and (ii)computation on computing nodes locally. The communication and computation are dealt with inside an operator and are not separable. In other words, it is difficult to avoid large intermediate results and hence reduce the communication cost. To reduce communication cost, we separate communication from computation using several new operators proposed in this paper. One is a pair operator () to pair the partitions of a relation R with the partitions of a relation S, where a partition is specified by a hash function. With the pair operator defined, we can explicitly deal with communication to deliver pairs of partitions to computing nodes. Together with , we can also explicitly treat the local computation on a computing node as op for any RA (relational algebra) operator op. We give a merge operator (U), to collect all partial results from computing nodes as they are. In short, with , op, and U, we are able to explicitly specify communication and computation for RA operators. Furthermore, we propose new techniques, namely, partitioning push-down and computation push-up to separate communication from computation for RA expressions. We prove that we can push-down/up for a wide range of relational expressions. We have developed a distributed system named Secco (Separate Communication from Computation) by revamping SparkSQL on Spark, and confirmed the efficiency of our approach in our performance studies using real datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Distributed Evaluation of Graph Queries Using Recursive Relational AlgebraSarah Chlyah, Pierre Genevès, Nabil LayaïdaICDE 2025
- New Query Optimization Techniques in the Spark Engine of Azure SynapseAbhishek Modi, Kaushik Rajan, Srinivas Thimmaiah, Prakhar Jain et al.VLDB 2022 · 13 citations
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
- Thrifty Query Execution via IncrementabilityDixin Tang, Zechao Shang, Aaron J. Elmore, Sanjay Krishnan et al.SIGMOD 2020 · 9 citations
- LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise ComparisonJunhao Ye, Jiahui Li, Lu Chen, Yuren Mao et al.VLDB 2025 · 2 citations
