TPCx-AI under the Microscope: A Benchmarking Debt Analysis
Ilin Tolovski, Philipp Hildebrandt, Khuzaima Daudjee, Tilmann Rabl
Abstract
TPCx-AI is an industry standard benchmark for evaluating the end-to-end performance of machine learning systems and the underlying hardware configurations. In the database community, individual parts of the dataset and the workloads are used to evaluate preprocessing methods and systems for fast inference. In both of these cases, the datasets and workloads are used based on the characteristics defined in the specification. Upon analysis of TPCx-AI's dataset and use cases, we observe that the official implementation of TPCx-AI's kit diverges from the specification, does not evaluate the capabilities of the system under test, and impacts the overall performance in a benchmark run.
In this paper, we investigate the benchmarking debt accumulated in the TPCx-AI dataset and the workloads. We identify properties that impact the benchmark's performance, including runtime and quality of use cases, the defined metrics and their thresholds, workload discrepancies, and data errors. Our analysis shows that all use cases and datasets contain benchmarking debts, impacting the training and serving runtimes by up to 350× and 800×, respectively. By addressing these debts, we observe an end-to-end throughput increase of up to 3.8× over the default TPCx-AI implementation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19feb39d-1f9e-49f0-982f-ffff719f1dd2Builds on7
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 91 citations
- Optimizing Data Pipelines for Machine Learning in Feature StoresRui Liu, Kwanghyun Park, Fotis Psallidas, Xiaoyong Zhu et al.VLDB 2023 · 10 citations
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 7 citations
Related papers
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 57 citations
- SQLStorm: Taking Database Benchmarking into the LLM EraTobias Schmidt, Viktor Leis, Peter Boncz, Thomas NeumannVLDB 2025 · 21 citations
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu et al.VLDB 2022 · 169 citations
- Toward Drift-Aware Database BenchmarkingGuanli Liu, Renata Borovica-GajicVLDB 2026
- DBPA: A Benchmark for Transactional Database Performance AnomaliesShiyue Huang, Ziwei Wang, Xinyi Zhang, Yaofeng Tu et al.SIGMOD 2023 · 8 citations
