Pre-training Summarization Models of Structured Datasets for Cardinality Estimation
Yao Lu, Srikanth Kandula, Arnd Christian König, Surajit Chaudhuri
摘要
We consider the problem of pre-training models which convert structured datasets into succinct summaries that can be used to answer cardinality estimation queries. Doing so avoids per-dataset training and, in our experiments, reduces the time to construct summaries by up to 100×. When datasets change, our summaries are incrementally updateable. Our key insights are to use multiple summaries per dataset, use learned summaries for columnsets for which other simpler techniques do not achieve high accuracy, and that analogous to similar pre-trained models for images and text, structured datasets have some common frequency and correlation patterns which our models learn to capture by pre-training on a large and diverse corpus of datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Kepler: Robust Learning for Parametric Query OptimizationLyric Doshi, Vincent Zhuang, Gaurav Jain, Ryan Marcus 等SIGMOD 2023 · 被引用 35 次
- Solving Max-Min Fair Resource Allocations Quickly on Large GraphsPooria Namyar, Behnaz Arzani, Srikanth Kandula, Santiago Segarra 等NSDI 2024 · 被引用 29 次
- PRICE: A Pretrained Model for Cross-Database Cardinality EstimationTianjing Zeng, Junwei Lan, Jiahong Ma, Wenqing Wei 等VLDB 2025 · 被引用 14 次
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha 等VLDB 2022 · 被引用 14 次
- AutoCE: An Accurate and Efficient Model Advisor for Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiICDE 2023 · 被引用 13 次
它引用的顶会 Paper4
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina 等VLDB 2020 · 被引用 154 次
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
- Approximate Partition Selection for Big-Data Workloads using Summary StatisticsKexin Rong, Yao Lu, Peter Bailis, Srikanth Kandula 等VLDB 2020 · 被引用 8 次
相关 Paper
- Cardinality Estimation of Approximate Substring Queries using Deep LearningSuyong Kwon, Woohwan Jung, Kyuseok ShimVLDB 2022 · 被引用 10 次
- Are We Ready For Learned Cardinality Estimation?Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang 等VLDB 2021 · 被引用 156 次
- BaCon: Efficient Batch Processing of Counting QueriesYuxi Liu, Xiao Hu, Pankaj K. Agarwal, Jun YangVLDB 2026
- Data-Agnostic Cardinality Learning from Imperfect WorkloadsPeizhi Wu, Rong Kang, Tieying Zhang, Jianjun Chen 等VLDB 2025 · 被引用 1 次
- Sample-Efficient Cardinality Estimation Using Geometric Deep LearningSilvan Reiner, Michael GrossniklausVLDB 2024 · 被引用 20 次
