LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration
Zhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham, Lin Ma
摘要
Lakehouse systems unify the strengths of data lakes and data warehouses and are rapidly becoming a dominant architecture for analytic data management. The lakehouse architecture decouples system design into interoperable subsystems—execution engines (e.g., Spark, Trino, Presto) and table formats (e.g., Delta Lake, Iceberg, Hudi)—giving users flexibility to mix and match. However, jointly selecting and configuring these subsystems is hard: subsystem choices and configurations interact in complex ways, and online trial-and-error is costly (or infeasible when migration is required). Although there is extensive work on database tuning, most methods target a single subsystem and thus miss cross-dependencies; many also rely on iterative online tuning that is prohibitively expensive. In this work, we present LakeHelm, a zero-shot lakehouse advisor that jointly recommends an engine-format pair and its configuration without online feedback. LakeHelm uses a dual-gate Mixture-of-Experts model: separate gates specialize in engine and format choices, and experts learn configuration surrogates for each subsystem combination. To enhance generalization, we augment training data with generated SQL templates and synthesized workloads, layered atop collected runs that explore the configuration space. Evaluated across five standard benchmarks (TPC-DS, TPC-H, JOB, SSB, SSB-Flat), LakeHelm delivers competitive execution times—averaging 1.35x speedup over a fixed overall-best lakehouse configuration across a large number of workload variations. It achieves this via zero-shot inference on unseen workloads in seconds, without costly online experimentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul 等SIGMOD 2021 · 被引用 242 次
- Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental EvaluationXinyi Zhang, Zhuo Chang, Yang Li, Hong Wu 等VLDB 2022 · 被引用 88 次
- LlamaTune: Sample-Efficient DBMS Configuration TuningKonstantinos Kanellis, Cong Ding, Brian Kroth, Andreas Müller 等VLDB 2022 · 被引用 73 次
- UDO: Universal Database Optimization using Reinforcement LearningJunxiong Wang, Immanuel Trummer, Debabrota BasuVLDB 2021 · 被引用 53 次
相关 Paper
- PTO: A Workload-driven Predictive Table Optimizer for Lakehouse SystemsVenkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold ReinwaldSIGMOD 2026
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia 等SIGMOD 2024 · 被引用 7 次
- Decisionhouse: Prescriptive Analytics in the Data StackMatteo Brucato, Fjodor Kholodkov, Soren Little, Jakob Mayer 等VLDB 2026 · 被引用 2 次
- TreeCat: Standalone Catalog Engine for Large Data SystemsKeonwoo Oh, Pooja Nilangekar, Amol DeshpandeVLDB 2025
- Interoperable ACID Transactions for Open Table FormatsTobias Götz, Daniel Ritter, Muhammad El-Hindi, Jana GicevaVLDB 2026
