PTO: A Workload-driven Predictive Table Optimizer for Lakehouse Systems
Venkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold Reinwald
摘要
Data lakehouse architectures manage both structured and semi-structured data, often using disaggregated storage with volumes that can reach petabyte scale of data stored in open table formats such as Apache Iceberg. Due to the size and storage structure, traditional indexes are cumbersome to maintain resulting in the need for effective table organization to enable efficient retrieval of relevant data for analytical queries. To maximize skipping of irrelevant data while scanning large tables, lakehouse systems rewrite data files according to pre-specified partitioning columns, target file sizes, row group sizes, and bin-packing or sort strategies. Optimizing these parameters can enhance the skipping of irrelevant data during table scans and improve query performance significantly. State-of-the-art lakehouse systems often require these parameters to be manually specified by the user which is impractical due to the combinatorial search space of parameter values thereby severely impeding the usability of the existing table optimization features in these systems. Conducting an exhaustive search to find the best combination of these parameters is impractical because these parameters are interdependent on each other, and rewriting a table with a single instantiation of all the four parameters can already take several hours at terabyte-scale. This comes with the additional complexity that optimal parameter value settings are query workload-sensitive as the filter predicates associated with the scan operators in the workload determine the skipping benefits we can get on a data layout. While our solution is applicable to lakehouse systems and open table formats which adopt similar parameterized layouts, we implemented PTO on Presto lakehouse engine to optimize Apache Iceberg tables. Our experiments show that PTO reduces the average workload latency by 11% on TPC-H and 36% on TPC-DS benchmarks at SF 10K while speeding up scan-intensive, long latency queries by 3.4× and 11× respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and ConfigurationZhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham 等VLDB 2026
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou 等VLDB 2023 · 被引用 10 次
- A Community Cache with Complete InformationMania Abdi, Amin Mosayyebzadeh, Mohammad Hossein Hajkazemi, Emine Ugur Kaynar 等FAST 2021 · 被引用 2 次
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia 等SIGMOD 2024 · 被引用 7 次
- TreeCat: Standalone Catalog Engine for Large Data SystemsKeonwoo Oh, Pooja Nilangekar, Amol DeshpandeVLDB 2025
