LAQy: Efficient and Reusable Query Approximations via Lazy Sampling
Viktor Sanca, Periklis Chrysogelos, Anastasia Ailamaki
摘要
Modern analytical engines rely on Approximate Query Processing (AQP) to provide faster response times than the hardware allows for exact query answering. However, existing AQP methods impose steep performance penalties as workload unpredictability increases. Specifically, offline AQP relies on predictable workloads to create samples that match the queries in a priori to query execution, reducing query response times when queries match the expected workload. As soon as workload predictability diminishes, existing online AQP methods create query-specific samples with little reuse across queries, producing significantly smaller gains in response times. As a result, existing approaches cannot fully exploit the benefits of sampling under increased unpredictability. We analyze sample creation and propose LAQy, a framework for building, expanding, and merging samples to adapt to the changes in workload predicates. We show the main parameters that affect the sample creation time and propose lazy sampling to overcome the unpredictability issues that cause fast-but-specialized samples to be query-specific. We evaluate LAQy by implementing it in an in-memory code-generation-based scale-up analytical engine to show the adaptivity and practicality of our framework in a modern system. LAQy speeds up online sampling processing as a function of sample reuse ranging from practically zero to full online sampling time and from 2.5x to 19.3x in a simulated exploratory workload.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- PairwiseHist: Fast, Accurate, and Space-Efficient Approximate Query Processing with Data CompressionAaron Hurst, Daniel E. Lucani, Qi ZhangVLDB 2024 · 被引用 5 次
- Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and QualityXilin Tang, Feng Zhang, Shuhao Zhang, Yani Liu 等SIGMOD 2025 · 被引用 1 次
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang 等SIGMOD 2025 · 被引用 1 次
相关 Paper
- An Agile Sample Maintenance Approach for Agile AnalyticsHanbing Zhang, Yazhong Zhang, Zhenying He, Yinan Jing 等ICDE 2020 · 被引用 2 次
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 被引用 54 次
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang 等VLDB 2023 · 被引用 9 次
- Permutable Compiled Queries: Dynamically Adapting Compiled Queries without RecompilingPrashanth Menon, Amadou Ngom, Todd C. Mowry, Andrew Pavlo 等VLDB 2021 · 被引用 26 次
- Thrifty Query Execution via IncrementabilityDixin Tang, Zechao Shang, Aaron J. Elmore, Sanjay Krishnan 等SIGMOD 2020 · 被引用 9 次
