Redbench: Workload Synthesis From Cloud Traces
Johannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid, Martin Stemmer, Andreas Kipf, Carsten Binnig, Muhammad El-Hindi
Abstract
Workload traces from cloud data warehouse providers reveal that standard benchmarks such as TPC-H and TPC-DS fail to capture key characteristics of real-world workloads, including query repetition and string-heavy queries. In this paper, we introduce Redbench, a novel benchmark featuring a workload generator that reproduces real-world workload characteristics derived from traces released by cloud providers. Redbench integrates multiple workload generation techniques to tailor workloads to specific objectives, transforming existing benchmarks into realistic query streams that preserve intrinsic workload characteristics. By focusing on inherent workload signals rather than execution-specific metrics, Redbench bridges the gap between synthetic and real workloads. Our evaluation shows that (1) Redbench produces more realistic and reproducible workloads for cloud data warehouse benchmarking, and (2) Redbench reveals the impact of system optimizations across four commercial data warehouse platforms. We believe that Redbench provides a crucial foundation for advancing research on optimization techniques for modern cloud data warehouses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ef4e1b2-86b0-4b0f-ab17-01119dc50ff6Cited by top-tier papers2
- Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database EnginesJohannes Wehrstein, Timo Eckmann, Matthias Jasny, Carsten BinnigVLDB 2026 · 17 citations
- Toward Drift-Aware Database BenchmarkingGuanli Liu, Renata Borovica-GajicVLDB 2026
Builds on10
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong et al.NSDI 2020 · 142 citations
- Flow-Loss: Learning Cardinality Estimates That MatterParimarjan Negi, Ryan Marcus, Andreas Kipf, Hongzi Mao et al.VLDB 2021 · 102 citations
- Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetAlexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya et al.VLDB 2024 · 65 citations
- Learning a Partitioning Advisor for Cloud DatabasesBenjamin Hilprecht, Carsten Binnig, Uwe RöhmSIGMOD 2020 · 64 citations
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 62 citations
Related papers
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang et al.VLDB 2025 · 4 citations
- Cloud Analytics BenchmarkAlexander van Renen, Viktor LeisVLDB 2023 · 32 citations
- HyBench: A New Benchmark for HTAP DatabasesChao Zhang, Guoliang Li, Tao LvVLDB 2024 · 28 citations
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 57 citations
- CloudyBench: A Testbed for A Comprehensive Evaluation of Cloud-Native DatabasesChao Zhang, Guoliang Li, Leyao Liu, Tao Lv et al.ICDE 2025 · 4 citations
