Learning to be a Statistician: Learned Estimator for Number of Distinct Values
Renzhi Wu, Bolin Ding, Xu Chu, Zhewei Wei, Xiening Dai, Tao Guan, Jingren Zhou
摘要
Estimating the number of distinct values (NDV) in a column is useful for many tasks in database systems, such as columnstore compression and data profiling. In this work, we focus on how to derive accurate NDV estimations from random (online/offline) samples. Such efficient estimation is critical for tasks where it is prohibitive to scan the data even once. Existing sample-based estimators typically rely on heuristics or assumptions and do not have robust performance across different datasets as the assumptions on data can easily break. On the other hand, deriving an estimator from a principled formulation such as maximum likelihood estimation is very challenging due to the complex structure of the formulation. We propose to formulate the NDV estimation task in a supervised learning framework, and aim to learn a model as the estimator. To this end, we need to answer several questions: i) how to make the learned model workload agnostic; ii) how to obtain training data; iii) how to perform model training. We derive conditions of the learning framework under which the learned model is workload agnostic , in the sense that the model/estimator can be trained with synthetically generated training data, and then deployed into any data warehouse simply as, e.g. , user-defined functions (UDFs), to offer efficient (within microseconds on CPU) and accurate NDV estimations for unseen tables and workloads. We compare the learned estimator with the state-of-the-art sample-based estimators on nine real-world datasets to demonstrate its superior estimation accuracy. We publish our code for training data generation, model training, and the learned estimator online for reproducibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- PilotScope: Steering Databases with Machine Learning DriversRong Zhu, Lianggui Weng, Wenqing Wei, Di Wu 等VLDB 2024 · 被引用 18 次
- Ground Truth Inference for Weakly Supervised Entity MatchingRenzhi Wu, Alexander Bendeck, Xu Chu, Yeye HeSIGMOD 2023 · 被引用 4 次
- Auctionformer: A Unified Deep Learning Algorithm for Solving Equilibrium Strategies in Auction GamesKexin Huang, Ziqian Chen, Xue Wang, Chongming Gao 等ICML 2024 · 被引用 3 次
- AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse EstimatorsXianghong Xu, Tieying Zhang, Xiao He, Haoyang Li 等VLDB 2025 · 被引用 3 次
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 被引用 2 次
它引用的顶会 Paper4
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du 等ICLR 2021 · 被引用 364 次
- Are We Ready For Learned Cardinality Estimation?Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang 等VLDB 2021 · 被引用 156 次
- FLAT: Fast, Lightweight and Accurate Method for Cardinality EstimationRong Zhu, Ziniu Wu, Yuxing Han, Kai Zeng 等VLDB 2021 · 被引用 120 次
- Efficiently Approximating Selectivity Functions using Low Overhead Regression ModelsAnshuman Dutt, Chi Wang, Vivek R. Narasayya, Surajit ChaudhuriVLDB 2020 · 被引用 45 次
相关 Paper
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang 等SIGMOD 2025 · 被引用 1 次
- Learning-based Property Estimation with PolynomialsJiajun Li, Runlin Lei, Sibo Wang, Zhewei Wei 等SIGMOD 2024 · 被引用 3 次
- Sampling-based Estimation of the Number of Distinct Values in Distributed EnvironmentJiajun Li, Zhewei Wei, Bolin Ding, Xiening Dai 等KDD 2022 · 被引用 5 次
- Sample-based Distinct Cardinality Estimation for Multiple Attributes in Multi-Dataset QueriesMehnaz Tabassum Mahin, Michael J. Carey, Vassilis J. TsotrasVLDB 2026
- Learned Cardinality Estimation: A Design Space Exploration and A Comparative EvaluationJi Sun, Jintao Zhang, Zhaoyan Sun, Guoliang Li 等VLDB 2022 · 被引用 90 次
