Grep: A Graph Learning Based Database Partitioning System
Xuanhe Zhou, Guoliang Li, Jianhua Feng, Luyang Liu, Wei Guo
摘要
Database partitioning is a fundamental but challenging task in distributed databases, which selects specific columns as a partitioning key for each table and uses the partitioning key to allocate the table data into different compute nodes in order to maximize the performance. However, this problem is NP-hard and existing distributed databases require users to manually specify the partitioning keys, which may cause potential performance degradation. Although reinforcement learning based methods have been proposed, they have several limitations. First, they do not capture the complex data distributions and query access patterns, and thus involve high computation cost across different compute nodes to answer a query. Second, they involve an expensive step to repetitively partition the data into different compute nodes in order to train a learned key-selection model, which is a waste of time and resources. To address these limitations, we propose a practical learned database partitioning system Grep. We first adopt a graph model to encode data and query features, where vertices are columns, edges are query relations, and the weights of columns are computed based on the localized graph structures (e.g., data diversity, joined columns). We then utilize graph neural networks to embed the partitioning factors into embedding vectors in order to capture the data and query correlations. Next we propose a key-selection model to select appropriate partitioning keys based on the graph model. Finally, we propose an evaluation model to estimate the partitioning performance without actually partitioning the database. We have implemented Grep in a commercial distributed database, and experiments show the effectiveness of our system (e.g., 68% higher throughput for 30K queries in a real banking scenario).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The Holon Approach for Simultaneously Tuning Multiple Components in a Self-Driving Database Management System with Machine Learning via Synthesized Proto-ActionsWilliam Zhang, Wan Shen Lim, Matthew Butrovich, Andrew PavloVLDB 2024 · 被引用 13 次
- OpenFGL: A Comprehensive Benchmark for Federated Graph LearningXunkai Li, Yinlin Zhu, Boyang Pang, Guochen Yan 等VLDB 2025 · 被引用 12 次
- PACE: Poisoning Attacks on Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiSIGMOD 2024 · 被引用 9 次
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 被引用 7 次
- Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP SystemsZhenghao Ding, Xinyi Zhang, Chao Zhang, Yishen Sun 等VLDB 2026 · 被引用 1 次
它引用的顶会 Paper13
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Reinforcement Learning with Tree-LSTM for Join Order SelectionXiang Yu, Guoliang Li, Chengliang Chai, Nan TangICDE 2020 · 被引用 168 次
- Query Performance Prediction for Concurrent Queries using Graph EmbeddingXuanhe Zhou, Ji Sun, Guoliang Li, Jianhua FengVLDB 2020 · 被引用 96 次
- Learned Cardinality Estimation: A Design Space Exploration and A Comparative EvaluationJi Sun, Jintao Zhang, Zhaoyan Sun, Guoliang Li 等VLDB 2022 · 被引用 90 次
- A Learned Query Rewrite System using Monte Carlo Tree SearchXuanhe Zhou, Guoliang Li, Chengliang Chai, Jianhua FengVLDB 2022 · 被引用 85 次
相关 Paper
- Adaptive Partitioning for Large-Scale Graph Analytics in Geo-Distributed Data CentersAmelie Chi Zhou, Juanyun Luo, Ruibo Qiu, Haobin Tan 等ICDE 2022 · 被引用 8 次
- Learning a Partitioning Advisor for Cloud DatabasesBenjamin Hilprecht, Carsten Binnig, Uwe RöhmSIGMOD 2020 · 被引用 64 次
- R2O: A Dual-Layer Framework for Joint Rewriting and Ordering in Distributed Property Graph Query OptimizationMin Shi, Peng Peng, Xin Xiao, Lei Zou 等SIGMOD 2026
- NeuroCut: A Neural Approach for Robust Graph PartitioningRishi Shah, Krishnanshu Jain, Sahil Manchanda, Sourav Medya 等KDD 2024 · 被引用 2 次
- PreQR: Pre-training Representation for SQL UnderstandingXiu Tang, Sai Wu, Mingli Song, Shanshan Ying 等SIGMOD 2022 · 被引用 21 次
