PTO: A Workload-driven Predictive Table Optimizer for Lakehouse Systems
Venkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold Reinwald
Abstract
Data lakehouse architectures manage both structured and semi-structured data, often using disaggregated storage with volumes that can reach petabyte scale of data stored in open table formats such as Apache Iceberg. Due to the size and storage structure, traditional indexes are cumbersome to maintain resulting in the need for effective table organization to enable efficient retrieval of relevant data for analytical queries. To maximize skipping of irrelevant data while scanning large tables, lakehouse systems rewrite data files according to pre-specified partitioning columns, target file sizes, row group sizes, and bin-packing or sort strategies. Optimizing these parameters can enhance the skipping of irrelevant data during table scans and improve query performance significantly. State-of-the-art lakehouse systems often require these parameters to be manually specified by the user which is impractical due to the combinatorial search space of parameter values thereby severely impeding the usability of the existing table optimization features in these systems. Conducting an exhaustive search to find the best combination of these parameters is impractical because these parameters are interdependent on each other, and rewriting a table with a single instantiation of all the four parameters can already take several hours at terabyte-scale. This comes with the additional complexity that optimal parameter value settings are query workload-sensitive as the filter predicates associated with the scan operators in the workload determine the skipping benefits we can get on a data layout. While our solution is applicable to lakehouse systems and open table formats which adopt similar parameterized layouts, we implemented PTO on Presto lakehouse engine to optimize Apache Iceberg tables. Our experiments show that PTO reduces the average workload latency by 11% on TPC-H and 36% on TPC-DS benchmarks at SF 10K while speeding up scan-intensive, long latency queries by 3.4× and 11× respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9cfdde1a-a4db-4bb5-b82a-c845fd114318Related papers
- LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and ConfigurationZhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham et al.VLDB 2026
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou et al.VLDB 2023 · 10 citations
- A Community Cache with Complete InformationMania Abdi, Amin Mosayyebzadeh, Mohammad Hossein Hajkazemi, Emine Ugur Kaynar et al.FAST 2021 · 2 citations
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia et al.SIGMOD 2024 · 7 citations
- TreeCat: Standalone Catalog Engine for Large Data SystemsKeonwoo Oh, Pooja Nilangekar, Amol DeshpandeVLDB 2025
