TPCx-AI under the Microscope: A Benchmarking Debt Analysis
Ilin Tolovski, Philipp Hildebrandt, Khuzaima Daudjee, Tilmann Rabl
摘要
TPCx-AI is an industry standard benchmark for evaluating the end-to-end performance of machine learning systems and the underlying hardware configurations. In the database community, individual parts of the dataset and the workloads are used to evaluate preprocessing methods and systems for fast inference. In both of these cases, the datasets and workloads are used based on the characteristics defined in the specification. Upon analysis of TPCx-AI's dataset and use cases, we observe that the official implementation of TPCx-AI's kit diverges from the specification, does not evaluate the capabilities of the system under test, and impacts the overall performance in a benchmark run.
In this paper, we investigate the benchmarking debt accumulated in the TPCx-AI dataset and the workloads. We identify properties that impact the benchmark's performance, including runtime and quality of use cases, the defined metrics and their thresholds, workload discrepancies, and data errors. Our analysis shows that all use cases and datasets contain benchmarking debts, impacting the training and serving runtimes by up to 350× and 800×, respectively. By addressing these debts, we observe an end-to-end throughput increase of up to 3.8× over the default TPCx-AI implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 被引用 91 次
- Optimizing Data Pipelines for Machine Learning in Feature StoresRui Liu, Kwanghyun Park, Fotis Psallidas, Xiaoyong Zhu 等VLDB 2023 · 被引用 10 次
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 被引用 7 次
相关 Paper
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 被引用 57 次
- SQLStorm: Taking Database Benchmarking into the LLM EraTobias Schmidt, Viktor Leis, Peter Boncz, Thomas NeumannVLDB 2025 · 被引用 21 次
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu 等VLDB 2022 · 被引用 169 次
- Toward Drift-Aware Database BenchmarkingGuanli Liu, Renata Borovica-GajicVLDB 2026
- DBPA: A Benchmark for Transactional Database Performance AnomaliesShiyue Huang, Ziwei Wang, Xinyi Zhang, Yaofeng Tu 等SIGMOD 2023 · 被引用 8 次
