Lune

ICDE2026顶会

Batcher: Learning to Construct Cost-Efficient Batches of Small Queries in Big Data Processing Platforms

Yeonsu Park, Taesung Lee, Byungchul Tak, Wook-Shin Han

2026年份

摘要

Diverse big data processing platforms play critical roles in modern data analytics systems. Their strengths lie in processing queries on huge volumes of data with high parallelism on distributed nodes. However, one type of workload, made of an excessive number of small queries, is known to pose performance issues by preventing big data processing platforms from reaching their intended performance. A recent technique of merging small queries into a large query mitigated this issue by enabling higher parallelism during query processing. However, we hypothesize that the methodical rearranging of queries into batches by similarity and submitting them, instead of one large single query, can produce significantly higher performance gains. To validate this, we have designed and implemented a system, called Batcher, that could learn the optimal batching strategies by utilizing a query batch cost estimation model and multi-staged cluster refinements to handle the NP-hard batch forming task with low overhead. Our evaluations of Batcher using a largescale real-world dataset showed that our strategy could achieve up to 5.4× improved performance.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖