Weighted Distinct Sampling: Cardinality Estimation for SPJ Queries
Yuan Qiu, Yilei Wang, Ke Yi, Feifei Li, Bin Wu, Chaoqun Zhan
Abstract
SPJ (select-project-join) queries form the backbone of many SQL queries used in practice. Accurate cardinality estimation of these queries is thus an important problem, with applications in query optimization, approximate query processing, and data analytics. However, this problem has not been rigorously addressed in the literature, despite the fact that cardinality estimation techniques of the three relational operators, selection, projection, and join, have each been extensively studied (but not when used in combination) in the past 30+ years. The major technical difficulty is that (distinct) projection seems to be difficult to combine with the other two operators when it comes to cardinality estimation.
In this paper, we give the first formal study of cardinality estimation for SP queries. While it was studied in a prior work in 2001, there is no guarantee on its optimality. We define a class of algorithms, which we call weighted distinct sampling, for estimating SP query sizes, and show how to find a near-optimal sampling strategy that is away from the optimum only by a lower order term. We then extend it to handling SPJ queries, giving the first non-trivial solution for SPJ cardinality estimation. We have also performed an extensive experimental evaluation to complement our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94d8b7c0-6ae3-4300-ba01-4f457458eed1Cited by top-tier papers6
- ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic WorkloadsPengfei Li, Wenqing Wei, Rong Zhu, Bolin Ding et al.VLDB 2024 · 50 citations
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha et al.VLDB 2022 · 14 citations
- Density-optimized Intersection-free Mapping and Matrix Multiplication for Join-Project OperationsZichun Huang, Shimin ChenVLDB 2022 · 9 citations
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 9 citations
- Towards a Converged Relational-Graph Optimization FrameworkYunkai Lou, Longbin Lai, Bingqing Lyu, Yufan Yang et al.SIGMOD 2025 · 4 citations
Related papers
- From Single to Multiple Attributes: Experimental Insights on Sampling-Based Distinct Combination Estimation in Group-by QueriesYujie Zhang, Xiaochun Yang, Bin Wang, Yuan SuiICDE 2026
- ASM: Harmonizing Autoregressive Model, Sampling, and Multi-dimensional Statistics Merging for Cardinality EstimationKyoungmin Kim, Sangoh Lee, Injung Kim, Wook-Shin HanSIGMOD 2024 · 18 citations
- Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset QueriesMehnaz Tabassum Mahin, Michael J. Carey, Vassilis J. TsotrasVLDB 2026
- LpBound: Pessimistic Cardinality Estimation Using ℓp-Norms of Degree SequencesHaozhe Zhang, Christoph Mayer, Mahmoud Abo Khamis, Dan Olteanu et al.SIGMOD 2025 · 7 citations
- A Practical Approach to Groupjoin and Nested AggregatesPhilipp Fent, Thomas NeumannVLDB 2021 · 11 citations
