LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration
Zhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham, Lin Ma
Abstract
Lakehouse systems unify the strengths of data lakes and data warehouses and are rapidly becoming a dominant architecture for analytic data management. The lakehouse architecture decouples system design into interoperable subsystems—execution engines (e.g., Spark, Trino, Presto) and table formats (e.g., Delta Lake, Iceberg, Hudi)—giving users flexibility to mix and match. However, jointly selecting and configuring these subsystems is hard: subsystem choices and configurations interact in complex ways, and online trial-and-error is costly (or infeasible when migration is required). Although there is extensive work on database tuning, most methods target a single subsystem and thus miss cross-dependencies; many also rely on iterative online tuning that is prohibitively expensive. In this work, we present LakeHelm, a zero-shot lakehouse advisor that jointly recommends an engine-format pair and its configuration without online feedback. LakeHelm uses a dual-gate Mixture-of-Experts model: separate gates specialize in engine and format choices, and experts learn configuration surrogates for each subsystem combination. To enhance generalization, we augment training data with generated SQL templates and synthesized workloads, layered atop collected runs that explore the configuration space. Evaluated across five standard benchmarks (TPC-DS, TPC-H, JOB, SSB, SSB-Flat), LakeHelm delivers competitive execution times—averaging 1.35x speedup over a fixed overall-best lakehouse configuration across a large number of workload variations. It achieves this via zero-shot inference on unseen workloads in seconds, without costly online experimentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f03d1e27-d817-45ff-9d44-eb01b303b885Builds on14
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental EvaluationXinyi Zhang, Zhuo Chang, Yang Li, Hong Wu et al.VLDB 2022 · 88 citations
- LlamaTune: Sample-Efficient DBMS Configuration TuningKonstantinos Kanellis, Cong Ding, Brian Kroth, Andreas Müller et al.VLDB 2022 · 73 citations
- UDO: Universal Database Optimization using Reinforcement LearningJunxiong Wang, Immanuel Trummer, Debabrota BasuVLDB 2021 · 53 citations
Related papers
- PTO: A Workload-driven Predictive Table Optimizer for Lakehouse SystemsVenkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold ReinwaldSIGMOD 2026
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia et al.SIGMOD 2024 · 7 citations
- Decisionhouse: Prescriptive Analytics in the Data StackMatteo Brucato, Fjodor Kholodkov, Soren Little, Jakob Mayer et al.VLDB 2026 · 2 citations
- TreeCat: Standalone Catalog Engine for Large Data SystemsKeonwoo Oh, Pooja Nilangekar, Amol DeshpandeVLDB 2025
- Interoperable ACID Transactions for Open Table FormatsTobias Götz, Daniel Ritter, Muhammad El-Hindi, Jana GicevaVLDB 2026
