IDEBench: A Benchmark for Interactive Data Exploration
Philipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim Kraska
摘要
Existing benchmarks for analytical database systems such as TPC-DS and TPC-H are designed for static reporting scenarios. The main metric of these benchmarks is the performance of running individual SQL queries over a synthetic database. In this paper, we argue that such benchmarks are not suitable for evaluating database workloads originating from interactive data exploration (IDE) systems where most queries are ad-hoc, not based on predefined reports, and built incrementally.
As a main contribution, we present a novel benchmark called IDEBench that can be used to evaluate the performance of database systems for IDE workloads. As opposed to traditional benchmarks for analytical database systems, our goal is to provide more meaningful workloads and datasets that can be used to benchmark IDE query engines, with a particular focus on metrics that capture the trade-off between query performance and quality of the result. As a second contribution, this paper evaluates and discusses the performance results of selected IDE query engines using our benchmark. The study includes two commercial systems, as well as two research prototypes (IDEA, approXimateD-B/XDB), and one traditional analytical database system (MonetDB).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina 等VLDB 2020 · 被引用 154 次
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai 等SIGMOD 2021 · 被引用 90 次
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 被引用 54 次
- Database Benchmarking for Supporting Real-Time Interactive Querying of Large DataLeilani Battle, Philipp Eichmann, Marco Angelini, Tiziana Catarci 等SIGMOD 2020 · 被引用 35 次
- Cache Me If You Can: Accuracy-Aware Inference Engine for Differentially Private Data ExplorationMiti Mazmudar, Thomas Humphries, Jiaxiang Liu, Matthew Rafuse 等VLDB 2023 · 被引用 15 次
相关 Paper
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 被引用 62 次
- An Adaptive Benchmark for Modeling User Exploration of Large DatasetsJoanna Purich, Anthony Wise, Leilani BattleSIGMOD 2025 · 被引用 1 次
- HyBench: A New Benchmark for HTAP DatabasesChao Zhang, Guoliang Li, Tao LvVLDB 2024 · 被引用 28 次
- Redbench: Workload Synthesis From Cloud TracesJohannes Wehrstein, Roman Heinrich, Mihail Stoian, Skander Krid 等VLDB 2026 · 被引用 7 次
- Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetAlexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya 等VLDB 2024 · 被引用 65 次
