A Step Toward Deep Online Aggregation
Nikhil Sheoran, Supawit Chockchowwat, Arav Chheda, Suwen Wang, Riya Verma, Yongjoo Park
Abstract
For exploratory data analysis, it is often desirable to know what answers you are likely to get before actually obtaining those answers. This can potentially be achieved by designing systems to offer the estimates of a data operation result-say op(data)-earlier in the process based on partial data processing. Those estimates continuously refine as more data is processed and finally converge to the exact answer. Unfortunately, the existing techniques-called Online Aggregation (OLA)-are limited to a single operation; that is, we cannot obtain the estimates for op(op(data)) or op(...(op(data))). If this Deep OLA becomes possible, data analysts will be able to explore data more interactively using complex cascade operations. In this work, we take a step toward Deep OLA with evolving data frames (edf), a novel data model to offer OLA for nested ops-op(...(op(data)))-by representing an evolving structured data (with converging estimates) that is closed under set operations. That is, op(edf) produces yet another edf; thus, we can freely apply successive operations to edf and obtain an OLA output for each op. We evaluate its viability with Wake, an edf-based OLA system, by examining against state-of-the-art OLA and non-OLA systems. In our experiments on TPC-H dataset, Wake produces its first estimates 4.93× faster (median)-with 1.3× median slowdown for exact answers-compared to conventional systems. Besides its generality, Wake is also 1.92× faster (median) than existing OLA systems in producing estimates of under 1% relative errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68471fb3-bf7d-4490-bd6a-5783d23a9eecCited by top-tier papers2
- S/C: Speeding up Data Materialization with Bounded MemoryZhaoheng Li, Xinyu Pi, Yongjoo ParkICDE 2023 · 7 citations
- Biathlon: Harnessing Model Resilience for Accelerating ML Inference PipelinesChaokun Chang, Eric Lo, Chunxiao YeVLDB 2024 · 5 citations
Builds on11
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- Are We Ready For Learned Cardinality Estimation?Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang et al.VLDB 2021 · 156 citations
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina et al.VLDB 2020 · 154 citations
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang et al.VLDB 2021 · 138 citations
- Fauce: Fast and Accurate Deep Ensembles with Uncertainty for Cardinality EstimationJie Liu, Wenqian Dong, Dong Li, Qingqing ZhouVLDB 2021 · 71 citations
Related papers
- EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized ViewsZhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, Umakishore RamachandranSIGMOD 2022 · 26 citations
- TEGRA: Efficient Ad-Hoc Analytics on Evolving GraphsAnand Padmanabha Iyer, Qifan Pu, Kishan Patel, Joseph E. Gonzalez et al.NSDI 2021
- Window Function Optimization: Co-Evaluation and Other TechniquesDaniel Lindner, Felix Naumann, Alberto LernerVLDB 2026
- Efficient GPU-Centric Evolving Graph Processing at ScaleYunmo Zhang, Jiacheng Huang, Xizhe Yin, Junqiao Qiu et al.OSDI 2026
- Rotary: A Resource Arbitration Framework for Progressive Iterative AnalyticsRui Liu, Aaron J. Elmore, Michael J. Franklin, Sanjay KrishnanICDE 2023 · 1 citation
