Dynamic Data Layout Optimization with Worst-Case Guarantees
Kexin Rong, Paul Liu, Sarah Ashok Sonje, Moses Charikar
摘要
Many data analytics systems store and process large datasets in partitions containing millions of rows. By mapping rows to partitions in an optimized way, it is possible to improve query performance by skipping over large numbers of irrelevant partitions during query processing. This mapping is referred to as a data layout. Recent works have shown that customizing the data layout to the anticipated query workload greatly improves query performance, but the performance benefits may disappear if the workload changes. Reorganizing data layouts to accommodate workload drift can resolve this issue, but reorganization costs could exceed query savings if not done carefully.
In this paper, we present an algorithmic framework OREO that makes online reorganization decisions to balance the benefits of improved query performance with the costs of reorganization. Our framework extends results from Metrical Task Systems to provide a tight bound on the worst-case performance guarantee for online reorganization, without prior knowledge of the query workload. Through evaluation on real-world datasets and query workloads, our experiments demonstrate that online reorganization with OREO can lead to an up to 32% improvement in combined query and reorganization time compared to using a single, optimized data layout for the entire workload.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HONEYBEE: Efficient Role-based Access Control for Vector Databases via Dynamic PartitioningHongbin Zhong, Matthew Lentz, Nina Narodytska, Adriana Szekeres 等SIGMOD 2026 · 被引用 5 次
- Breaking the Isolation-Freshness Trade-off: Joint Adaptive Storage Optimization for HTAP SystemsZhenghao Ding, Xinyi Zhang, Chao Zhang, Yishen Sun 等VLDB 2026 · 被引用 1 次
它引用的顶会 Paper11
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 被引用 180 次
- LISA: A Learned Index Structure for Spatial DataPengfei Li, Hua Lu, Qian Zheng, Long Yang 等SIGMOD 2020 · 被引用 158 次
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke 等SIGMOD 2020 · 被引用 87 次
- Learning a Partitioning Advisor for Cloud DatabasesBenjamin Hilprecht, Carsten Binnig, Uwe RöhmSIGMOD 2020 · 被引用 64 次
- OPTIMUSCLOUD: Heterogeneous Configuration Optimization for Distributed Databases in the CloudAshraf Mahgoub, Alexander Medoff, Rakesh Kumar, Subrata Mitra 等USENIX ATC 2020 · 被引用 63 次
相关 Paper
- Workload-Aware Incremental Reclustering in Cloud Data WarehousesYipeng Liu, Renfei Zhou, Jiaqi Yan, Huanchen ZhangSIGMOD 2026 · 被引用 1 次
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang 等SIGMOD 2021 · 被引用 37 次
- Pando: Enhanced Data Skipping with Logical Data PartitioningSivaprasad Sudhir, Wenbo Tao, Nikolay Pavlovich Laptev, Cyrille Habis 等VLDB 2023 · 被引用 14 次
- PTO: A Workload-driven Predictive Table Optimizer for Lakehouse SystemsVenkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold ReinwaldSIGMOD 2026
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha 等ICDE 2021 · 被引用 15 次
