Efficient Approximate Algorithms for Empirical Variance with Hashed Block Sampling
Xingguang Chen, Fangyuan Zhang, Sibo Wang
Abstract
Empirical variance is a fundamental concept widely used in data management and data analytics, e.g., query optimization, approximate query processing, and feature selection. A direct solution to derive the empirical variance is scanning the whole data table, which is expensive when the data size is huge. Hence, most current works focus on approximate answers by sampling. For results with approximation guarantees, the samples usually need to be uniformly independent random, incurring high cache miss rates especially in compact columnar style layouts. An alternative uses block sampling to avoid this issue, which directly samples a block of consecutive records fitting page sizes instead of sampling one record each time. However, this provides no theoretical guarantee. Existing studies show that the practical estimations can be inaccurate as the records within a block can be correlated.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2ba85855-2ce0-4ced-b3de-73d53e12b2a3Cited by top-tier papers2
- PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error GuaranteesYuxuan Zhu, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang et al.SIGMOD 2025 · 3 citations
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang et al.SIGMOD 2025 · 1 citation
Related papers
- Random Sampling for Group-By QueriesTrong Duc Nguyen, Ming-Hung Shih, Sai Sree Parvathaneni, Bojian Xu et al.ICDE 2020 · 12 citations
- Approximate Partition Selection for Big-Data Workloads using Summary StatisticsKexin Rong, Yao Lu, Peter Bailis, Srikanth Kandula et al.VLDB 2020 · 8 citations
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang et al.VLDB 2023 · 9 citations
- Sharp Empirical Bernstein Inequalities for the Variance of Bounded Random VariablesDiego Martinez Taboada, Aaditya RamdasICML 2026 · 6 citations
- Instance-Optimality in I/O-Efficient Sampling and Sequential EstimationShyam Narayanan, Václav Rozhon, Jakub Tetek, Mikkel ThorupFOCS 2024
