Datamime: Generating Representative Benchmarks by Automatically Synthesizing Datasets
Hyun Ryong Lee, Daniel Sánchez
摘要
Benchmarks that closely match the behavior of production workloads are crucial to design and provision computer systems. However, current approaches fall short: First, open-source benchmarks use public datasets that cause different behavior from production workloads. Second, blackbox workload cloning techniques generate synthetic code that imitates the target workload, but the resulting program fails to capture most workload characteristics, such as microarchitectural bottlenecks or time-varying behavior.
Generating code that mimics a complex application is an extremely hard problem. Instead, we propose a different and easier approach to benchmark synthesis. Our key insight is that, for many production workloads, the program is publicly available or there is a reasonably similar open-source program. In this case, generating the right dataset is sufficient to produce an accurate benchmark.
Based on this observation, we present Datamime, a profileguided approach to generate representative benchmarks for production workloads. Datamime uses the performance profiles of a target workload to generate a dataset that, when used by a benchmark program, behaves very similarly to the target workload in terms of its microarchitectural characteristics.
We evaluate Datamime on several datacenter workloads. Datamime generates synthetic benchmarks that closely match the microarchitectural features of these workloads, with a mean absolute percentage error of 3.2% on IPC. Microarchitectural behavior stays close across processor types. Finally, timevarying behaviors are also replicated, making these benchmarks useful to e.g. characterize and optimize tail latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- SATORI: Efficient and Fair Resource Partitioning by Sacrificing Short-Term Benefits for Long-Term Gains*Rohan Basu Roy, Tirthak Patel, Devesh TiwariISCA 2021 · 被引用 34 次
相关 Paper
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang 等VLDB 2025 · 被引用 4 次
- Redbench: Workload Synthesis From Cloud TracesJohannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid 等VLDB 2026 · 被引用 7 次
- Ditto: End-to-End Application Cloning for Networked Cloud ServicesMingyu Liang, Yu Gan, Yueying Li, Carlos Torres 等ASPLOS 2023 · 被引用 10 次
- SAM: Database Generation from Query Workloads with Supervised Autoregressive ModelsJingyi Yang, Peizhi Wu, Gao Cong, Tieying Zhang 等SIGMOD 2022 · 被引用 13 次
- Mystique: Enabling Accurate and Scalable Generation of Production AI BenchmarksMingyu Liang, Wenyin Fu, Louis Feng, Zhongyi Lin 等ISCA 2023 · 被引用 9 次
