Batcher: Learning to Construct Cost-Efficient Batches of Small Queries in Big Data Processing Platforms
Yeonsu Park, Taesung Lee, Byungchul Tak, Wook-Shin Han
Abstract
Diverse big data processing platforms play critical roles in modern data analytics systems. Their strengths lie in processing queries on huge volumes of data with high parallelism on distributed nodes. However, one type of workload, made of an excessive number of small queries, is known to pose performance issues by preventing big data processing platforms from reaching their intended performance. A recent technique of merging small queries into a large query mitigated this issue by enabling higher parallelism during query processing. However, we hypothesize that the methodical rearranging of queries into batches by similarity and submitting them, instead of one large single query, can produce significantly higher performance gains. To validate this, we have designed and implemented a system, called Batcher, that could learn the optimal batching strategies by utilizing a query batch cost estimation model and multi-staged cluster refinements to handle the NP-hard batch forming task with low overhead. Our evaluations of Batcher using a largescale real-world dataset showed that our strategy could achieve up to 5.4× improved performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0a5d73fe-7303-4e79-a244-d197f49ec5c4Related papers
- QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in SparkYeonsu Park, Byungchul Tak, Wook-Shin HanSIGMOD 2023 · 3 citations
- Hitcher: Efficient GPU-based Vector Search via Cluster-Centric Kernel and Hitch-Ride OrderingQihui Zhou, Changji Li, Guanxian Jiang, Chenhao Ma et al.KDD 2026
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 76 citations
- Adaptive Code Generation for Data-Intensive AnalyticsWangda Zhang, Junyoung Kim, Kenneth A. Ross, Eric Sedlar et al.VLDB 2021 · 12 citations
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
