QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark
Yeonsu Park, Byungchul Tak, Wook-Shin Han
Abstract
Spark big data processing platform is heavily used in today's IT services for various critical applications such as machine learning tasks for service recommendations or massive volumes of raw sales data analysis. Spark is designed to deliver high performance by enabling a high degree of parallelism while processing various heavy-weight queries that require homogeneous operations on large data. However, it has been observed that workloads made of small and short-running queries coming from various sources are becoming dominant in practice. Unfortunately, the current Spark architecture is unfit to process workloads made of a large number of small queries optimally due to excessive I/Os with small computations. We present a technique, called QaaD, that addresses this problem fundamentally by applying i) transparent conversion of workloads made of small queries into one with large queries and ii) dynamic partition size adjustment for runtime overhead minimization. For this, we introduce a new abstraction, microRDD, to support our design of query merging, the embedding of queries as part of data, and an opportunistic sharing of common input data among queries. Comprehensive evaluation using real-world data shows that QaaD is able to deliver 10.6x to 36.6x speed-up against standard Spark executions for small query workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aa7df355-80cb-40cf-8b20-eee106dfcda3Related papers
- SASPAR: Shared Adaptive Stream PartitioningJeyhun Karimov, Hans-Arno JacobsenICDE 2023 · 3 citations
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 9 citations
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha et al.ICDE 2021 · 15 citations
- Batcher: Learning to Construct Cost-Efficient Batches of Small Queries in Big Data Processing PlatformsYeonsu Park, Taesung Lee, Byungchul Tak, Wook-Shin HanICDE 2026
- Conjunctive Queries with ComparisonsQichen Wang, Ke YiSIGMOD 2022 · 13 citations
