Check Out the Big Brain on BRAD: Simplifying Cloud Data Processing with Learned Automated Data Meshes
Tim Kraska, Tianyu Li, Samuel Madden, Markos Markakis, Amadou Ngom, Ziniu Wu, Geoffrey X. Yu
摘要
The last decade of database research has led to the prevalence of specialized systems for different workloads. Consequently, organizations often rely on a combination of specialized systems, organized in a Data Mesh. Data meshes present significant challenges for system administrators, including picking the right system for each workload, moving data between systems, maintaining consistency, and correctly configuring each system. Many non-expert end users (e.g., data analysts or app developers) either cannot solve their business problems, or suffer from sub-optimal performance or cost due to this complexity. We envision BRAD, a cloud system that automatically integrates and manages data and systems into an instance-optimized data mesh, allowing users to efficiently store and query data under a unified data model (i.e., relational tables) without knowledge of underlying system details. With machine learning, BRAD automatically deduces the strengths and weaknesses of each engine through a combination of offline training and online probing. Then, BRAD uses these insights to route queries to the most suitable (combination of) system(s) for efficient execution. Furthermore, BRAD automates configuration tuning, resource scaling, and data migration across component systems, and makes recommendations for more impactful decisions, such as adding or removing systems. As such, BRAD exemplifies a new class of systems that utilize machine learning and the cloud to make complex data processing more accessible to end users, raising numerous new problems in database systems, machine learning, and the cloud.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Blueprinting the Cloud: Unifying and Automatically Optimizing Cloud Data Infrastructures with BRADGeoffrey X. Yu, Ziniu Wu, Ferdinand Kossmann, Tianyu Li 等VLDB 2024 · 被引用 11 次
- Fast and Scalable Data Transfer Across Data SystemsHaralampos Gavriilidis, Kaustubh Beedkar, Matthias Boehm, Volker MarklSIGMOD 2025 · 被引用 4 次
- Improving DBMS Scheduling Decisions with Accurate Performance Prediction on Concurrent QueriesZiniu Wu, Markos Markakis, Chunwei Liu, Peter Baile Chen 等VLDB 2025
- A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume MetricGunika Verma, Aashutosh A V, Pooja Srinivas, Yogesh Simmhan 等VLDB 2026
它引用的顶会 Paper18
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul 等SIGMOD 2021 · 被引用 242 次
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 被引用 180 次
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 被引用 178 次
- Reinforcement Learning with Tree-LSTM for Join Order SelectionXiang Yu, Guoliang Li, Chengliang Chai, Nan TangICDE 2020 · 被引用 168 次
相关 Paper
- SageDB: An Instance-Optimized Data Analytics SystemJialin Ding, Ryan Marcus, Andreas Kipf, Vikram Nathan 等VLDB 2022 · 被引用 18 次
- Declarative Data Serving: The Future of Machine Learning Inference on the EdgeTed Shaowang, Nilesh Jain, Dennis Matthews, Sanjay KrishnanVLDB 2021 · 被引用 15 次
- OPTIMUSCLOUD: Heterogeneous Configuration Optimization for Distributed Databases in the CloudAshraf Mahgoub, Alexander Medoff, Rakesh Kumar, Subrata Mitra 等USENIX ATC 2020 · 被引用 63 次
- Towards Dynamic and Safe Configuration Tuning for Cloud DatabasesXinyi Zhang, Hong Wu, Yang Li, Jian Tan 等SIGMOD 2022 · 被引用 62 次
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 被引用 34 次
