LAQy: Efficient and Reusable Query Approximations via Lazy Sampling
Viktor Sanca, Periklis Chrysogelos, Anastasia Ailamaki
Abstract
Modern analytical engines rely on Approximate Query Processing (AQP) to provide faster response times than the hardware allows for exact query answering. However, existing AQP methods impose steep performance penalties as workload unpredictability increases. Specifically, offline AQP relies on predictable workloads to create samples that match the queries in a priori to query execution, reducing query response times when queries match the expected workload. As soon as workload predictability diminishes, existing online AQP methods create query-specific samples with little reuse across queries, producing significantly smaller gains in response times. As a result, existing approaches cannot fully exploit the benefits of sampling under increased unpredictability. We analyze sample creation and propose LAQy, a framework for building, expanding, and merging samples to adapt to the changes in workload predicates. We show the main parameters that affect the sample creation time and propose lazy sampling to overcome the unpredictability issues that cause fast-but-specialized samples to be query-specific. We evaluate LAQy by implementing it in an in-memory code-generation-based scale-up analytical engine to show the adaptivity and practicality of our framework in a modern system. LAQy speeds up online sampling processing as a function of sample reuse ranging from practically zero to full online sampling time and from 2.5x to 19.3x in a simulated exploratory workload.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d8bbf5cc-cb1e-4d17-a8c8-92a73ada81a5Cited by top-tier papers3
- PairwiseHist: Fast, Accurate, and Space-Efficient Approximate Query Processing with Data CompressionAaron Hurst, Daniel E. Lucani, Qi ZhangVLDB 2024 · 5 citations
- Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and QualityXilin Tang, Feng Zhang, Shuhao Zhang, Yani Liu et al.SIGMOD 2025 · 1 citation
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang et al.SIGMOD 2025 · 1 citation
Related papers
- An Agile Sample Maintenance Approach for Agile AnalyticsHanbing Zhang, Yazhong Zhang, Zhenying He, Yinan Jing et al.ICDE 2020 · 2 citations
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 54 citations
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang et al.VLDB 2023 · 9 citations
- Permutable Compiled Queries: Dynamically Adapting Compiled Queries without RecompilingPrashanth Menon, Amadou Ngom, Todd C. Mowry, Andrew Pavlo et al.VLDB 2021 · 26 citations
- Thrifty Query Execution via IncrementabilityDixin Tang, Zechao Shang, Aaron J. Elmore, Sanjay Krishnan et al.SIGMOD 2020 · 9 citations
