Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database Engines
Johannes Wehrstein, Timo Eckmann, Matthias Jasny, Carsten Binnig
Abstract
Modern OLAP engines support arbitrary analytical workloads, but this flexibility incurs overhead from runtime schema interpretation, generic data representations, and abstraction layers, even in compiled-query systems. Workload-specific engines can eliminate these costs and exploit specialized data structures and algorithms for higher performance, yet have historically been too expensive to build manually. Recent advances in LLM-based code synthesis challenge this tradeoff, but naive prompting does not produce correct or efficient engines due to deep architectural dependencies and the need for systematic refinement. We present Bespoke OLAP , a fully autonomous synthesis pipeline that constructs high-performance OLAP engines tailored to a target workload through iterative performance evaluation and automated validation. Bespoke OLAP generates engines from scratch within minutes to hours and achieves order-of-magnitude speedups over DuckDB and Umbra, demonstrating that the generality tax extends beyond query compilation to storage layout and algorithmic design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dcd046a-0132-4677-803d-ed64fb9d1d71Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Flow-Loss: Learning Cardinality Estimates That MatterParimarjan Negi, Ryan Marcus, Andreas Kipf, Hongzi Mao et al.VLDB 2021 · 102 citations
- CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 CodexImmanuel TrummerVLDB 2022 · 77 citations
- Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetAlexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya et al.VLDB 2024 · 65 citations
- Redbench: Workload Synthesis From Cloud TracesJohannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid et al.VLDB 2026 · 7 citations
- GRACEFUL: A Learned Cost Estimator for UDFsJohannes Wehrstein, Tiemo Bang, Roman Heinrich, Carsten BinnigICDE 2025 · 2 citations
Related papers
- Charting the Design Space of Query Execution using VOILATim Gubner, Peter BonczVLDB 2021 · 17 citations
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang et al.VLDB 2026 · 5 citations
- Incremental Fusion: Unifying Compiled and Vectorized Query ExecutionBenjamin Wagner, André Kohn, Peter Boncz, Viktor LeisICDE 2024 · 3 citations
- Automating Database-Native Function Code Synthesis with LLMsWei Zhou, Xuanhe Zhou, Qikang He, Guoliang Li et al.SIGMOD 2026 · 6 citations
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 7 citations
