Blueprinting the Cloud: Unifying and Automatically Optimizing Cloud Data Infrastructures with BRAD
Geoffrey X. Yu, Ziniu Wu, Ferdinand Kossmann, Tianyu Li, Markos Markakis, Amadou Latyr Ngom, Samuel Madden, Tim Kraska
摘要
Modern organizations manage their data with a wide variety of specialized cloud database engines (e.g., Aurora, BigQuery, etc.). However, designing and managing such infrastructures is hard. Developers must consider many possible designs with non-obvious performance consequences; moreover, current software abstractions tightly couple applications to specific systems (e.g., with engine-specific clients), making it difficult to change after initial deployment. A better solution would virtualize cloud data management, allowing developers to declaratively specify their workload requirements and rely on automated solutions to design and manage the physical realization. In this paper, we present a technique called blueprint planning that achieves this vision. The key idea is to project data infrastructure design decisions into a unified design space (blueprints). We then systematically search over candidate blueprints using cost-based optimization, leveraging learned models to predict the utility of a blueprint on the workload. We use this technique to build BRAD, the first cloud data virtualization system. BRAD users issue queries to a single SQL interface that can be backed by multiple cloud database services. BRAD automatically selects the most suitable engine for each query, provisions and manages resources to minimize costs, and evolves the infrastructure to adapt to workload shifts. Our evaluation shows that BRAD meet user-defined performance targets and improve cost-savings by 1.6--13× compared to serverless auto-scaling or HTAP systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL WorkloadsJiale Lao, Immanuel TrummerSIGMOD 2026 · 被引用 11 次
- T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision TreesMaximilian Rieger, Thomas NeumannSIGMOD 2025 · 被引用 5 次
- CloudGlide: Deconstructing the Landscape of Cloud-Based AnalyticsMichail Georgoulakis Misegiannis, Daniel Ritter, Viktor Leis, Jana GicevaVLDB 2025 · 被引用 1 次
- Improving DBMS Scheduling Decisions with Accurate Performance Prediction on Concurrent QueriesZiniu Wu, Markos Markakis, Chunwei Liu, Peter Baile Chen 等VLDB 2025
- AQD: Online Adaptive Query Dispatcher for HTAP DatabasesYang Wu, Tongliang Li, Xuanhe Zhou, Jianying Wang 等VLDB 2026
它引用的顶会 Paper25
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul 等SIGMOD 2021 · 被引用 242 次
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 被引用 180 次
相关 Paper
- Check Out the Big Brain on BRAD: Simplifying Cloud Data Processing with Learned Automated Data MeshesTim Kraska, Tianyu Li, Samuel Madden, Markos Markakis 等VLDB 2023 · 被引用 13 次
- Charting the Design Space of Query Execution using VOILATim Gubner, Peter BonczVLDB 2021 · 被引用 17 次
- Scaling a Declarative Cluster Manager Architecture with Query Optimization TechniquesKexin Rong, Mihai Budiu, Athinagoras Skiadopoulos, Lalith Suresh 等VLDB 2023 · 被引用 3 次
- Excalibur: A Virtual Machine for Adaptive Fine-grained JIT-Compiled Query Execution based on VOILATim Gubner, Peter BonczVLDB 2023 · 被引用 10 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
