PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
Yan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang, Magnus Mueller, Chao Zhang, Songyue Zhang, Pascal Pfeil, Dominik Horn, Zhengchun Liu, Davide Pagano, Tim Kraska
摘要
Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture the real execution statistics of production cloud workloads. Although some cloud database vendors have recently released real workload traces, these traces alone do not qualify as benchmarks, as they typically lack essential components like the original SQL queries and their underlying databases. To overcome this limitation, this paper introduces a new problem of workload synthesis with real statistics, which aims to generate synthetic workloads that closely approximate real execution statistics, including key performance metrics and operator distributions, in real cloud workloads. To address this problem, we propose PBench, a novel workload synthesizer that constructs synthetic workloads by judiciously selecting and combining workload components (i.e., queries and databases) from existing benchmarks. This paper studies the key challenges in PBench. First, we address the challenge of balancing performance metrics and operator distributions by introducing a multi-objective optimization-based component selection method. Second, to capture the temporal dynamics of real workloads, we design a timestamp assignment method that progressively refines workload timestamps. Third, to handle the disparity between the original workload and the candidate workload, we propose a component augmentation approach that leverages large language models (LLMs) to generate additional workload components while maintaining statistical fidelity. We evaluate PBench on real cloud workload traces, demonstrating that it reduces approximation error by up to 6× compared to state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Redbench: Workload Synthesis From Cloud TracesJohannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid 等VLDB 2026 · 被引用 7 次
- Mimesys: Generating Realistic Executable Testing Environments from Resource Usage TracesDonghyun Kim, Zichao Hu, Joydeep Biswas, Aditya Akella 等OSDI 2026
- LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and ConfigurationZhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham 等VLDB 2026
它引用的顶会 Paper8
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong 等NSDI 2020 · 被引用 142 次
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 被引用 91 次
- Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetAlexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya 等VLDB 2024 · 被引用 65 次
- LLM-R2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query EfficiencyZhaodonghui Li, Haitao Yuan, Huiming Wang, Gao Cong 等VLDB 2025 · 被引用 52 次
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 被引用 34 次
相关 Paper
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL WorkloadsJiale Lao, Immanuel TrummerSIGMOD 2026 · 被引用 11 次
- HyBench: A New Benchmark for HTAP DatabasesChao Zhang, Guoliang Li, Tao LvVLDB 2024 · 被引用 28 次
- SQLStorm: Taking Database Benchmarking into the LLM EraTobias Schmidt, Viktor Leis, Peter Boncz, Thomas NeumannVLDB 2025 · 被引用 21 次
- Cloud Analytics BenchmarkAlexander van Renen, Viktor LeisVLDB 2023 · 被引用 32 次
- CloudyBench: A Testbed for A Comprehensive Evaluation of Cloud-Native DatabasesChao Zhang, Guoliang Li, Leyao Liu, Tao Lv 等ICDE 2025 · 被引用 4 次
