Data Chunk Compaction in Vectorized Execution
Yiming Qiao, Huanchen Zhang
Abstract
Modern analytical database management systems often adopt vectorized query execution engines that process columnar data in batches (i.e., data chunks) to minimize the interpretation overhead and improve CPU parallelism. However, certain database operators, especially hash joins, can drastically reduce the number of valid entries in a data chunk, resulting in numerous small chunks in an execution pipeline. These small chunks cannot fully enjoy the benefits of vectorized query execution, causing significant performance degradation. The key research question is when and how to compact these small data chunks during query execution. In this paper, we first model the chunk compaction problem and analyze the trade-offs between different compaction strategies. We then propose a learning-based algorithm that can adjust the compaction threshold dynamically at run time. To answer the ''how'' question, we propose a compaction method for the hash join operator, called logical compaction, that minimizes data movements when compacting data chunks. We implemented the proposed techniques in the state-of-the-art DuckDB and observed up to 63% speedup when evaluated using the Join Order Benchmark, TPC-H, and TPC-DS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2521bfd6-d0b3-4e3a-8acf-2f18ebbd2b3aCited by top-tier papers1
Ask how each one uses itBuilds on6
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- To Partition, or Not to Partition, That is the Join Question in a Real SystemMaximilian Bandle, Jana Giceva, Thomas NeumannSIGMOD 2021 · 43 citations
- Charting the Design Space of Query Execution using VOILATim Gubner, Peter BonczVLDB 2021 · 17 citations
- Efficient Query Re-optimization with Judicious Subquery SelectionsJunyi Zhao, Huanchen Zhang, Yihan GaoSIGMOD 2023 · 12 citations
- Analyzing Vectorized Hash Tables Across CPU ArchitecturesMaximilian Böther, Lawrence Benson, Ana Klimovic, Tilmann RablVLDB 2023 · 11 citations
Related papers
- Saving Private Hash JoinLaurens Kuiper, Paul Gross, Peter Boncz, Hannes MühleisenVLDB 2025
- Selective Late Materialization in Modern Analytical DatabasesYihao Liu, Shaoxuan Tang, Yulong Hui, Hangrui Zhou et al.VLDB 2025
- These Rows Are Made for Sorting and That's Just What We'll DoLaurens Kuiper, Hannes MühleisenICDE 2023 · 6 citations
- Debunking the Myth of Join Ordering: Toward Robust SQL AnalyticsJunyi Zhao, Kai Su, Yifei Yang, Xiangyao Yu et al.SIGMOD 2025 · 13 citations
- Robust External Hash Aggregation in the Solid State AgeLaurens Kuiper, Peter Boncz, Hannes MühleisenICDE 2024 · 7 citations
