Efficient Query Processing with Optimistically Compressed Hash Tables & Strings in the USSR
Tim Gubner, Viktor Leis, Peter Boncz
Abstract
Modern query engines rely heavily on hash tables for query processing. Overall query performance and memory footprint is often determined by how hash tables and the tuples within them are represented. In this work, we propose three complementary techniques to improve this representation: Domain-Guided Prefix Suppression bit-packs keys and values tightly to reduce hash table record width. Optimistic Splitting decomposes values (and operations on them) into (operations on) frequently-accessed and infrequently-accessed value slices. By removing the infrequently-accessed value slices from the hash table record, it improves cache locality. The Unique Strings Self-aligned Region (USSR) accelerates handling frequently-occurring strings, which are very common in real-world data sets, by creating an on-the-fly dictionary of the most frequent strings. This allows executing many string operations with integer logic and reduces memory pressure. We integrated these techniques into Vectorwise. On the TPC-H benchmark, our approach reduces peak memory consumption by 2-4× and improves performance by up to 1.5×. On a real-world BI workload, we measured a 2× improvement in performance and in micro-benchmarks we observed speedups of up to 25×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a8d2b79-8c3e-4b55-b19a-2276e6f04234Cited by top-tier papers1
Ask how each one uses itRelated papers
- RABIT: Efficient Range Queries with Bitmap IndexingJunchang Wang, Fu Xiao, Manos AthanassoulisSIGMOD 2026
- Data Chunk Compaction in Vectorized ExecutionYiming Qiao, Huanchen ZhangSIGMOD 2025 · 2 citations
- PIDS: Attribute Decomposition for Improved Compression and Query Performance in Columnar StorageHao Jiang, Chunwei Liu, Qi Jin, John Paparrizos et al.VLDB 2020
- Analyzing Vectorized Hash Tables Across CPU ArchitecturesMaximilian Böther, Lawrence Benson, Ana Klimovic, Tilmann RablVLDB 2023 · 11 citations
- Order-Preserving Key Compression for In-Memory Search TreesHuanchen Zhang, Xiaoxuan Liu, David G. Andersen, Michael Kaminsky et al.SIGMOD 2020 · 31 citations
