AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators
Xianghong Xu, Tieying Zhang, Xiao He, Haoyang Li, Rong Kang, Wang Shuai, Linhui Xu, Zhimin Liang, Shangyu Luo, Lei Zhang, Jianjun Chen
Abstract
Estimating the Number of Distinct Values (NDV) is fundamental for numerous data management tasks, especially within database applications. However, most existing works primarily focus on introducing new statistical or learned estimators, while identifying the most suitable estimator for a given scenario remains largely unexplored. Therefore, we propose AdaNDV, a learned method designed to adaptively select and fuse existing estimators to address this issue. Specifically, (1) we propose to use learned models to distinguish between overestimated and underestimated estimators and then select appropriate estimators from each category. This strategy provides a complementary perspective by integrating overestimations and underestimations for error correction, thereby improving the accuracy of NDV estimation. (2) To further integrate the estimation results, we introduce a novel fusion approach that employs a learned model to predict the weights of the selected estimators and then applies a weighted sum to merge them. By combining these strategies, the proposed AdaNDV fundamentally distinguishes itself from previous works that directly estimate NDV. Moreover, extensive experiments conducted on real-world datasets, with the number of individual columns being several orders of magnitude larger than in previous studies, demonstrate the superior performance of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d99f2354-c36d-4f40-8689-1988e25d7daaCited by top-tier papers2
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang et al.SIGMOD 2025 · 1 citation
- From Single to Multiple Attributes: Experimental Insights on Sampling-Based Distinct Combination Estimation in Group-by QueriesYujie Zhang, Xiaochun Yang, Bin Wang, Yuan SuiICDE 2026
Builds on9
- ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic WorkloadsPengfei Li, Wenqing Wei, Rong Zhu, Bolin Ding et al.VLDB 2024 · 50 citations
- GitTables: A Large-Scale Corpus of Relational TablesMadelon Hulsebos, Çagatay Demiralp, Paul GrothSIGMOD 2023 · 42 citations
- An Efficient Transfer Learning Based Configuration Adviser for Database TuningXinyi Zhang, Hong Wu, Yang Li, Zhengju Tang et al.VLDB 2024 · 25 citations
- Learning to be a Statistician: Learned Estimator for Number of Distinct ValuesRenzhi Wu, Bolin Ding, Xu Chu, Zhewei Wei et al.VLDB 2022 · 16 citations
- AutoCE: An Accurate and Efficient Model Advisor for Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiICDE 2023 · 13 citations
Related papers
- Sampling-based Estimation of the Number of Distinct Values in Distributed EnvironmentJiajun Li, Zhewei Wei, Bolin Ding, Xiening Dai et al.KDD 2022 · 5 citations
- Learning-based Property Estimation with PolynomialsJiajun Li, Runlin Lei, Sibo Wang, Zhewei Wei et al.SIGMOD 2024 · 3 citations
- Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset QueriesMehnaz Tabassum Mahin, Michael J. Carey, Vassilis J. TsotrasVLDB 2026
- ASM: Harmonizing Autoregressive Model, Sampling, and Multi-dimensional Statistics Merging for Cardinality EstimationKyoungmin Kim, Sangoh Lee, Injung Kim, Wook-Shin HanSIGMOD 2024 · 18 citations
- ACE: A Cardinality Estimator for Set-Valued QueriesYufan Sheng, Xin Cao, Kaiqi Zhao, Yixiang Fang et al.VLDB 2025
