Data-Semantics-Aware Recommendation of Diverse Pivot Tables
Whanhee Cho, Anna Fariha
Abstract
Data summarization is essential to discover insights from large datasets. In spreadsheets, pivot tables offer a convenient way to summarize tabular data by computing aggregates over some attributes, grouped by others. However, identifying attribute combinations that will result in useful pivot tables remains a challenge, especially for high-dimensional datasets. We formalize the problem of automatically recommending insightful and interpretable pivot tables, eliminating the tedious manual process. A crucial aspect of recommending a set of pivot tables is to diversify them. Traditional work inadequately address the table-diversification problem, which leads us to the problem of pivot table diversification . We present SAGE, a data- s emantics- a ware system for recommendin g k-budgeted diverse pivot tables, overcoming the shortcomings of prior work for top-k recommendations that cause redundancy. SAGE ensures that each pivot table is insightful , interpretable , and adaptive to the user's actions and preferences, while also guaranteeing that the set of pivot tables are different from each other, offering a diverse recommendation. We make two key technical contributions: (1) a data-semantics-aware model to measure the utility of a single pivot table and the diversity of a set of pivot tables, and (2) a scalable greedy algorithm that can efficiently select a set of diverse pivot tables of high utility, by leveraging data semantics to significantly reduce the combinatorial search space. Our extensive experiments on four real-world datasets show that SAGE outperforms alternative approaches, and efficiently scales to accommodate high-dimensional datasets. Additionally, through multiple case studies, we demonstrate SAGE's qualitative superiority over existing tools, and through a user study, we validate its practical usefulness and alignment with user preferences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
- DashBot: Insight-Driven Dashboard Generation Based on Deep Reinforcement LearningDazhen Deng, Aoyu Wu, Huamin Qu, Yingcai WuIEEE VIS 2022 · 40 citations
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 22 citations
- Conformance Constraint Discovery: Measuring Trust in Data-Driven SystemsAnna Fariha, Ashish Tiwari, Arjun Radhakrishna, Sumit Gulwani et al.SIGMOD 2021 · 18 citations
Related papers
- Table2Analysis: Modeling and Recommendation of Common Analysis Patterns for Multi-Dimensional DataMengyu Zhou, Wang Tao, Pengxin Ji, Han Shi et al.AAAI 2020 · 26 citations
- FEDEX: An Explainability Framework for Data Exploration StepsDaniel Deutch, Amir Gilad, Tova Milo, Amit Mualem et al.VLDB 2022 · 15 citations
- Selecting Sub-tables for Data ExplorationYael Amsterdamer, Susan B. Davidson, Tova Milo, Kathy Razmadze et al.ICDE 2023 · 5 citations
- LakeVisage: Towards Scalable, Flexible and Interactive Visualization Recommendation for Data Discovery over Data LakesYihao Hu, Jin Wang, Sajjadur RahmanVLDB 2025 · 3 citations
- Tab-Shapley: Identifying Top-k Tabular Data Quality InsightsManisha Padala, Lokesh Nagalapatti, Atharv Tyagi, Ramasuri Narayanam et al.AAAI 2025 · 1 citation
