AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators
Xianghong Xu, Tieying Zhang, Xiao He, Haoyang Li, Rong Kang, Wang Shuai, Linhui Xu, Zhimin Liang, Shangyu Luo, Lei Zhang, Jianjun Chen
摘要
Estimating the Number of Distinct Values (NDV) is fundamental for numerous data management tasks, especially within database applications. However, most existing works primarily focus on introducing new statistical or learned estimators, while identifying the most suitable estimator for a given scenario remains largely unexplored. Therefore, we propose AdaNDV, a learned method designed to adaptively select and fuse existing estimators to address this issue. Specifically, (1) we propose to use learned models to distinguish between overestimated and underestimated estimators and then select appropriate estimators from each category. This strategy provides a complementary perspective by integrating overestimations and underestimations for error correction, thereby improving the accuracy of NDV estimation. (2) To further integrate the estimation results, we introduce a novel fusion approach that employs a learned model to predict the weights of the selected estimators and then applies a weighted sum to merge them. By combining these strategies, the proposed AdaNDV fundamentally distinguishes itself from previous works that directly estimate NDV. Moreover, extensive experiments conducted on real-world datasets, with the number of individual columns being several orders of magnitude larger than in previous studies, demonstrate the superior performance of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang 等SIGMOD 2025 · 被引用 1 次
- From Single to Multiple Attributes: Experimental Insights on Sampling-Based Distinct Combination Estimation in Group-by QueriesYujie Zhang, Xiaochun Yang, Bin Wang, Yuan SuiICDE 2026
它引用的顶会 Paper9
- ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic WorkloadsPengfei Li, Wenqing Wei, Rong Zhu, Bolin Ding 等VLDB 2024 · 被引用 50 次
- GitTables: A Large-Scale Corpus of Relational TablesMadelon Hulsebos, Çagatay Demiralp, Paul GrothSIGMOD 2023 · 被引用 42 次
- An Efficient Transfer Learning Based Configuration Adviser for Database TuningXinyi Zhang, Hong Wu, Yang Li, Zhengju Tang 等VLDB 2024 · 被引用 25 次
- Learning to be a Statistician: Learned Estimator for Number of Distinct ValuesRenzhi Wu, Bolin Ding, Xu Chu, Zhewei Wei 等VLDB 2022 · 被引用 16 次
- AutoCE: An Accurate and Efficient Model Advisor for Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiICDE 2023 · 被引用 13 次
相关 Paper
- Sampling-based Estimation of the Number of Distinct Values in Distributed EnvironmentJiajun Li, Zhewei Wei, Bolin Ding, Xiening Dai 等KDD 2022 · 被引用 5 次
- Learning-based Property Estimation with PolynomialsJiajun Li, Runlin Lei, Sibo Wang, Zhewei Wei 等SIGMOD 2024 · 被引用 3 次
- Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset QueriesMehnaz Tabassum Mahin, Michael J. Carey, Vassilis J. TsotrasVLDB 2026
- ASM: Harmonizing Autoregressive Model, Sampling, and Multi-dimensional Statistics Merging for Cardinality EstimationKyoungmin Kim, Sangoh Lee, Injung Kim, Wook-Shin HanSIGMOD 2024 · 被引用 18 次
- ACE: A Cardinality Estimator for Set-Valued QueriesYufan Sheng, Xin Cao, Kaiqi Zhao, Yixiang Fang 等VLDB 2025
