Data-Semantics-Aware Recommendation of Diverse Pivot Tables
Whanhee Cho, Anna Fariha
摘要
Data summarization is essential to discover insights from large datasets. In spreadsheets, pivot tables offer a convenient way to summarize tabular data by computing aggregates over some attributes, grouped by others. However, identifying attribute combinations that will result in useful pivot tables remains a challenge, especially for high-dimensional datasets. We formalize the problem of automatically recommending insightful and interpretable pivot tables, eliminating the tedious manual process. A crucial aspect of recommending a set of pivot tables is to diversify them. Traditional work inadequately address the table-diversification problem, which leads us to the problem of pivot table diversification . We present SAGE, a data- s emantics- a ware system for recommendin g k-budgeted diverse pivot tables, overcoming the shortcomings of prior work for top-k recommendations that cause redundancy. SAGE ensures that each pivot table is insightful , interpretable , and adaptive to the user's actions and preferences, while also guaranteeing that the set of pivot tables are different from each other, offering a diverse recommendation. We make two key technical contributions: (1) a data-semantics-aware model to measure the utility of a single pivot table and the diversity of a set of pivot tables, and (2) a scalable greedy algorithm that can efficiently select a set of diverse pivot tables of high utility, by leveraging data semantics to significantly reduce the combinatorial search space. Our extensive experiments on four real-world datasets show that SAGE outperforms alternative approaches, and efficiently scales to accommodate high-dimensional datasets. Additionally, through multiple case studies, we demonstrate SAGE's qualitative superiority over existing tools, and through a user study, we validate its practical usefulness and alignment with user preferences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
- DashBot: Insight-Driven Dashboard Generation Based on Deep Reinforcement LearningDazhen Deng, Aoyu Wu, Huamin Qu, Yingcai WuIEEE VIS 2022 · 被引用 40 次
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 被引用 22 次
- Conformance Constraint Discovery: Measuring Trust in Data-Driven SystemsAnna Fariha, Ashish Tiwari, Arjun Radhakrishna, Sumit Gulwani 等SIGMOD 2021 · 被引用 18 次
相关 Paper
- Table2Analysis: Modeling and Recommendation of Common Analysis Patterns for Multi-Dimensional DataMengyu Zhou, Wang Tao, Pengxin Ji, Han Shi 等AAAI 2020 · 被引用 26 次
- FEDEX: An Explainability Framework for Data Exploration StepsDaniel Deutch, Amir Gilad, Tova Milo, Amit Mualem 等VLDB 2022 · 被引用 15 次
- Selecting Sub-tables for Data ExplorationYael Amsterdamer, Susan B. Davidson, Tova Milo, Kathy Razmadze 等ICDE 2023 · 被引用 5 次
- LakeVisage: Towards Scalable, Flexible and Interactive Visualization Recommendation for Data Discovery over Data LakesYihao Hu, Jin Wang, Sajjadur RahmanVLDB 2025 · 被引用 3 次
- Tab-Shapley: Identifying Top-k Tabular Data Quality InsightsManisha Padala, Lokesh Nagalapatti, Atharv Tyagi, Ramasuri Narayanam 等AAAI 2025 · 被引用 1 次
