Blueprinting the Cloud: Unifying and Automatically Optimizing Cloud Data Infrastructures with BRAD
Geoffrey X. Yu, Ziniu Wu, Ferdinand Kossmann, Tianyu Li, Markos Markakis, Amadou Latyr Ngom, Samuel Madden, Tim Kraska
Abstract
Modern organizations manage their data with a wide variety of specialized cloud database engines (e.g., Aurora, BigQuery, etc.). However, designing and managing such infrastructures is hard. Developers must consider many possible designs with non-obvious performance consequences; moreover, current software abstractions tightly couple applications to specific systems (e.g., with engine-specific clients), making it difficult to change after initial deployment. A better solution would virtualize cloud data management, allowing developers to declaratively specify their workload requirements and rely on automated solutions to design and manage the physical realization. In this paper, we present a technique called blueprint planning that achieves this vision. The key idea is to project data infrastructure design decisions into a unified design space (blueprints). We then systematically search over candidate blueprints using cost-based optimization, leveraging learned models to predict the utility of a blueprint on the workload. We use this technique to build BRAD, the first cloud data virtualization system. BRAD users issue queries to a single SQL interface that can be backed by multiple cloud database services. BRAD automatically selects the most suitable engine for each query, provisions and manages resources to minimize costs, and evolves the infrastructure to adapt to workload shifts. Our evaluation shows that BRAD meet user-defined performance targets and improve cost-savings by 1.6--13× compared to serverless auto-scaling or HTAP systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 897a7bd6-795e-493f-9dc0-5e29ab4a2acaCited by top-tier papers5
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL WorkloadsJiale Lao, Immanuel TrummerSIGMOD 2026 · 11 citations
- T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision TreesMaximilian Rieger, Thomas NeumannSIGMOD 2025 · 5 citations
- CloudGlide: Deconstructing the Landscape of Cloud-Based AnalyticsMichail Georgoulakis Misegiannis, Daniel Ritter, Viktor Leis, Jana GicevaVLDB 2025 · 1 citation
- Improving DBMS Scheduling Decisions with Accurate Performance Prediction on Concurrent QueriesZiniu Wu, Markos Markakis, Chunwei Liu, Peter Baile Chen et al.VLDB 2025
- AQD: Online Adaptive Query Dispatcher for HTAP DatabasesYang Wu, Tongliang Li, Xuanhe Zhou, Jianying Wang et al.VLDB 2026
Builds on25
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 180 citations
Related papers
- Check Out the Big Brain on BRAD: Simplifying Cloud Data Processing with Learned Automated Data MeshesTim Kraska, Tianyu Li, Samuel Madden, Markos Markakis et al.VLDB 2023 · 13 citations
- Charting the Design Space of Query Execution using VOILATim Gubner, Peter BonczVLDB 2021 · 17 citations
- Scaling a Declarative Cluster Manager Architecture with Query Optimization TechniquesKexin Rong, Mihai Budiu, Athinagoras Skiadopoulos, Lalith Suresh et al.VLDB 2023 · 3 citations
- Excalibur: A Virtual Machine for Adaptive Fine-grained JIT-Compiled Query Execution based on VOILATim Gubner, Peter BonczVLDB 2023 · 10 citations
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
