Cloud Analytics Benchmark
Alexander van Renen, Viktor Leis
摘要
The cloud facilitates the transition to a service-oriented perspective. This affects cloud-native data management in general, and data analytics in particular. Instead of managing a multi-node database cluster on-premise, end users simply send queries to a managed cloud data warehouse and receive results. While this is obviously very attractive for end users, database system architects still have to engineer systems for this new service model. There are currently many competing architectures ranging from self-hosted (Presto, PostgreSQL), over managed (Snowflake, Amazon Redshift) to query-as-a-service (Amazon Athena, Google BigQuery) offerings. Benchmarking these architectural approaches is currently difficult, and it is not even clear what the metrics for a comparison should be. To overcome these challenges, we first analyze a real-world query trace from Snowflake and compare its properties to that of TPC-H and TPC-DS. Doing so, we identify important differences that distinguish traditional benchmarks from real-world cloud data warehouse workloads. Based on this analysis, we propose the Cloud Analytics Benchmark (CAB). By incorporating workload fluctuations and multi-tenancy, CAB allows evaluating different designs in terms of user-centered metrics such as cost and performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT PipelinesTengjun Jin, Yuxuan Zhu, Daniel KangVLDB 2026 · 被引用 13 次
- How Good are Learned Cost Models, Really? Insights from Query Optimization TasksRoman Heinrich, Manisha Luthra, Johannes Wehrstein, Harald Kornmayer 等SIGMOD 2025 · 被引用 13 次
- High-Performance Query Processing with NVMe Arrays: Spilling without Killing PerformanceMaximilian Kuschewski, Jana Giceva, Thomas Neumann, Viktor LeisSIGMOD 2025 · 被引用 11 次
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL WorkloadsJiale Lao, Immanuel TrummerSIGMOD 2026 · 被引用 11 次
它引用的顶会 Paper3
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong 等NSDI 2020 · 被引用 142 次
- Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud InfrastructureIngo Müller, Renato Marroquín, Gustavo AlonsoSIGMOD 2020 · 被引用 135 次
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 被引用 34 次
相关 Paper
- Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetAlexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya 等VLDB 2024 · 被引用 65 次
- CloudyBench: A Testbed for A Comprehensive Evaluation of Cloud-Native DatabasesChao Zhang, Guoliang Li, Leyao Liu, Tao Lv 等ICDE 2025 · 被引用 4 次
- Redbench: Workload Synthesis From Cloud TracesJohannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid 等VLDB 2026 · 被引用 7 次
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang 等VLDB 2025 · 被引用 4 次
- CloudGlide: Deconstructing the Landscape of Cloud-Based AnalyticsMichail Georgoulakis Misegiannis, Daniel Ritter, Viktor Leis, Jana GicevaVLDB 2025 · 被引用 1 次
