A Practical Approach to Groupjoin and Nested Aggregates
Philipp Fent, Thomas Neumann
Abstract
Groupjoins, the combined execution of a join and a subsequent group by, are common in analytical queries, and occur in about 1/8 of the queries in TPC-H and TPC-DS. While they were originally invented to improve performance, efficient parallel execution of groupjoins can be limited by contention, which limits their usefulness in a many-core system. Having an efficient implementation of groupjoins is highly desirable, as groupjoins are not only used to fuse group by and join but are also introduced by the unnesting component of the query optimizer to avoid nested-loops evaluation of aggregates. Furthermore, the query optimizer needs be able to reason over the result of aggregation in order to schedule it correctly. Traditional selectivity and cardinality estimations quickly reach their limits when faced with computed columns from nested aggregates, which leads to poor cost estimations and thus, suboptimal query plans.
In this paper, we present techniques to efficiently estimate, plan, and execute groupjoins and nested aggregates. We propose two novel techniques, aggregate estimates to predict the result distribution of aggregates, and parallel groupjoin execution for a scalable execution of groupjoins. The resulting system has significantly better estimates and a contention-free evaluation of groupjoins, which can speed up some TPC-H queries up to a factor of 2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ea7332e-18ee-4f9f-b497-17afa339c142Cited by top-tier papers5
- Database Technology for the Masses: Sub-Operators as First-Class EntitiesMaximilian Bandle, Jana GicevaVLDB 2021 · 19 citations
- A Scalable and Generic Approach to Range JoinsMaximilian Reif, Thomas NeumannVLDB 2022 · 6 citations
- Incremental Fusion: Unifying Compiled and Vectorized Query ExecutionBenjamin Wagner, André Kohn, Peter Boncz, Viktor LeisICDE 2024 · 3 citations
- Global Hash Tables Strike Back! An Analysis of Parallel GROUP BY AggregationDaniel Xue, Ryan MarcusVLDB 2026
- Cardinality Estimation for Having-ClausesGuido MoerkotteVLDB 2025
Builds on4
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 91 citations
- To Partition, or Not to Partition, That is the Join Question in a Real SystemMaximilian Bandle, Jana Giceva, Thomas NeumannSIGMOD 2021 · 43 citations
- Building Advanced SQL Analytics From Low-Level Plan OperatorsAndré Kohn, Viktor Leis, Thomas NeumannSIGMOD 2021 · 13 citations
- Getting Swole: Generating Access-Aware Code with Predicate PullupsAndrew Crotty, Alex Galakatos, Tim KraskaICDE 2020 · 11 citations
Related papers
- From Single to Multiple Attributes: Experimental Insights on Sampling-Based Distinct Combination Estimation in Group-by QueriesYujie Zhang, Xiaochun Yang, Bin Wang, Yuan SuiICDE 2026
- ASM: Harmonizing Autoregressive Model, Sampling, and Multi-dimensional Statistics Merging for Cardinality EstimationKyoungmin Kim, Sangoh Lee, Injung Kim, Wook-Shin HanSIGMOD 2024 · 18 citations
- Avoiding Materialisation for Guarded Aggregate QueriesMatthias Lanzinger, Reinhard Pichler, Alexander SelzerVLDB 2025 · 7 citations
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi et al.OSDI 2020 · 5 citations
- Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL ServerKukjin Lee, Anshuman Dutt, Vivek R. Narasayya, Surajit ChaudhuriVLDB 2023 · 26 citations
