PairwiseHist: Fast, Accurate, and Space-Efficient Approximate Query Processing with Data Compression
Aaron Hurst, Daniel E. Lucani, Qi Zhang
Abstract
Exponential growth in data collection is creating significant challenges for data storage and analytics latency. Approximate Query Processing (AQP) has long been touted as a solution for accelerating analytics on large datasets, however, there is still room for improvement across all key performance criteria. In this paper, we propose a novel histogram-based data synopsis called PairwiseHist that uses recursive hypothesis testing to ensure accurate histograms and can be built on top of data compressed using Generalized Deduplication (GD). We thus show that GD data compression can contribute to AQP. Compared to state-of-the-art AQP approaches, Pairwise-Hist achieves better performance across all key metrics, including 2.6× higher accuracy, 3.5× lower latency, 24× smaller synopses and 1.5--4× faster construction time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6e69339-99ca-4794-872d-1c3326352c8aCited by top-tier papers2
- SemBench: A Benchmark for Semantic Query Processing EnginesJiale Lao, Andreas Zimmerer, Olga Ovcharenko, Tianji Cong et al.VLDB 2026 · 31 citations
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian et al.VLDB 2025 · 2 citations
Builds on8
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina et al.VLDB 2020 · 154 citations
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 57 citations
- A Randomly Accessible Lossless Compression Scheme for Time-Series DataRasmus Vestergaard, Daniel E. Lucani, Qi ZhangINFOCOM 2020 · 29 citations
- Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query ProcessingXi Liang, Stavros Sintos, Zechao Shang, Sanjay KrishnanSIGMOD 2021 · 27 citations
- Conditional Generative Model Based Predicate-Aware Query ApproximationNikhil Sheoran, Subrata Mitra, Vibhor Porwal, Siddharth Ghetia et al.AAAI 2022 · 14 citations
Related papers
- Optimizing Random Access to Hierarchically-Compressed Data on GPUFeng Zhang, Yihua Hu, Haipeng Ding, Zhiming Yao et al.SC 2022 · 5 citations
- JanusAQP: Efficient Partition Tree Maintenance for Dynamic Approximate Query ProcessingXi Liang, Stavros Sintos, Sanjay KrishnanICDE 2023 · 3 citations
- LHist: Towards Learning Multi-dimensional Histogram for Massive Spatial DataQiyu Liu, Yanyan Shen, Lei ChenICDE 2021 · 20 citations
- DeepSketch: A New Machine Learning-Based Reference Search Technique for Post-Deduplication Delta CompressionJisung Park, Jeonggyun Kim, Yeseong Kim, Sungjin Lee et al.FAST 2022 · 40 citations
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 54 citations
