Exploiting Cloud Object Storage for High-Performance Analytics
Dominik Durner, Viktor Leis, Thomas Neumann
摘要
Elasticity of compute and storage is crucial for analytical cloud database systems. All cloud vendors provide disaggregated object stores, which can be used as storage backend for analytical query engines. Until recently, local storage was unavoidable to process large tables efficiently due to the bandwidth limitations of the network infrastructure in public clouds. However, the gap between remote network and local NVMe bandwidth is closing, making cloud storage more attractive. This paper presents a blueprint for performing efficient analytics directly on cloud object stores. We derive cost- and performance-optimal retrieval configurations for cloud object stores with the first in-depth study of this foundational service in the context of analytical query processing. For achieving high retrieval performance, we present AnyBlob , a novel download manager for query engines that optimizes throughput while minimizing CPU usage. We discuss the integration of high-performance data retrieval in query engines and demonstrate it by incorporating AnyBlob in our database system Umbra. Our experiments show that even without caching, Umbra with integrated AnyBlob achieves similar performance to state-of-the-art cloud data warehouses that cache data on local SSDs while improving resource elasticity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Cloud-Native Database Systems and Unikernels: Reimagining OS Abstractions for Modern HardwareViktor Leis, Christian DietrichVLDB 2024 · 被引用 12 次
- Two Birds With One Stone: Designing a Hybrid Cloud Storage Engine for HTAPTobias Schmidt, Dominik Durner, Viktor Leis, Thomas NeumannVLDB 2024 · 被引用 12 次
- High-Performance Query Processing with NVMe Arrays: Spilling without Killing PerformanceMaximilian Kuschewski, Jana Giceva, Thomas Neumann, Viktor LeisSIGMOD 2025 · 被引用 11 次
- F3: The Open-Source Data File Format for the FutureXinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney 等SIGMOD 2026 · 被引用 10 次
- CXL Memory Performance for In-Memory Data ProcessingMarcel Weisgut, Daniel Ritter, Pinar Tözün, Lawrence Benson 等VLDB 2025 · 被引用 7 次
它引用的顶会 Paper7
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong 等NSDI 2020 · 被引用 142 次
- Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud InfrastructureIngo Müller, Renato Marroquín, Gustavo AlonsoSIGMOD 2020 · 被引用 135 次
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu 等VLDB 2021 · 被引用 67 次
- The Case for Distributed Shared-Memory Databases with RDMA-Enabled Memory DisaggregationRuihong Wang, Jianguo Wang, Stratos Idreos, M. Tamer Özsu 等VLDB 2023 · 被引用 49 次
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 被引用 34 次
相关 Paper
- Crystal: A Unified Cache Storage System for Analytical DatabasesDominik Durner, Badrish Chandramouli, Yinan LiVLDB 2021 · 被引用 11 次
- Fusion: An Analytics Object Store Optimized for Query PushdownJianan Lu, Ashwini Raina, Asaf Cidon, Michael J. FreedmanASPLOS 2025 · 被引用 4 次
- Building Advanced SQL Analytics From Low-Level Plan OperatorsAndré Kohn, Viktor Leis, Thomas NeumannSIGMOD 2021 · 被引用 13 次
- Memory-Optimized Multi-Version Concurrency Control for Disk-Based Database SystemsMichael J. Freitag, Alfons Kemper, Thomas NeumannVLDB 2022 · 被引用 14 次
- CaaS-LSM: Compaction-as-a-Service for LSM-based Key-Value Stores in Storage Disaggregated InfrastructureQiaolin Yu, Chang Guo, Jay Zhuang, Viraj Thakkar 等SIGMOD 2024 · 被引用 18 次
