Lightweight Materialization for Fast Dashboards Over Joins
Zezhou Huang, Eugene Wu
摘要
Dashboards are vital in modern business intelligence tools, providing non-technical users with an interface to access comprehensive business data. With the rise of cloud technology, there is an increased number of data sources to provide enriched contexts for various analytical tasks, leading to a demand for interactive dashboards over a large number of joins. Nevertheless, joins are among the most expensive operations in DBMSes, making the support of interactive dashboards over joins challenging. In this paper, we present Treant, a dashboard accelerator for queries over large joins. Treant uses factorized query execution to handle aggregation queries over large joins, which alone is still insufficient for interactive speeds. To address this, we exploit the incremental nature of user interactions using Calibrated Junction Hypertree (CJT), a novel data structure that applies lightweight materialization of the intermediates during factorized execution. CJT ensures that the work needed to compute a query is proportional to how different it is from the previous query, rather than the overall complexity. Treant manages CJTs to share work between queries and performs materialization offline or during user "think-times." Implemented as a middleware that rewrites SQL, Treant is portable to any SQL-based DBMS. Our experiments on single node and cloud DBMSes show that Treant improves dashboard interactions by two orders of magnitude, and provides 10x improvement for ML augmentation compared to SOTA factorized ML system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 被引用 57 次
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez 等VLDB 2020 · 被引用 14 次
- Data Market Platforms: Trading Data Assets to Solve Data ProblemsRaul Castro Fernandez, Pranav Subramaniam, Michael J. FranklinVLDB 2020
相关 Paper
- Efficient Incrementialization of Correlated Nested Aggregate Queries using Relative Partial Aggregate Indexes (RPAI)Supun Abeysinghe, Qiyang He, Tiark RompfSIGMOD 2022 · 被引用 7 次
- Distributed Numerical and Machine Learning Computations via Two-Phase Execution of Aggregated Join TreesDimitrije Jankov, Binhang Yuan, Shangyu Luo, Chris JermaineVLDB 2021 · 被引用 9 次
- OM3: An Ordered Multi-level Min-Max Representation for Interactive Progressive Visualization of Time SeriesYunhai Wang, Yuchun Wang, Xin Chen, Yue Zhao 等SIGMOD 2023 · 被引用 8 次
- A Practical Approach to Groupjoin and Nested AggregatesPhilipp Fent, Thomas NeumannVLDB 2021 · 被引用 11 次
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu 等SIGMOD 2026 · 被引用 2 次
