Guided Exploration of Data Summaries
Brit Youngmann, Sihem Amer-Yahia, Aurélien Personnaz
Abstract
Data summarization is the process of producing interpretable and representative subsets of an input dataset. It is usually performed following a one-shot process with the purpose of finding the best summary. A useful summary contains k individually uniform sets that are collectively diverse to be representative. Uniformity addresses interpretability and diversity addresses representativity. Finding such as summary is a difficult task when data is highly diverse and large. We examine the applicability of Exploratory Data Analysis (EDA) to data summarization and formalize Eda4Sum, the problem of guided exploration of data summaries that seeks to sequentially produce connected summaries with the goal of maximizing their cumulative utility. Eda4Sum generalizes one-shot summarization. We propose to solve it with one of two approaches: (i) Top1Sum that chooses the most useful summary at each step; (ii) RLSum that trains a policy with Deep Reinforcement Learning that rewards an agent for finding a diverse and new collection of uniform sets at each step. We compare these approaches with one-shot summarization and top-performing EDA solutions. We run extensive experiments on three large datasets. Our results demonstrate the superiority of our approaches for summarizing very large data, and the need to provide guidance to domain experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc6e384d-6a1f-4bc6-af00-f892d9cb25edCited by top-tier papers7
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 12 citations
- Finding Convincing Views to Endorse a ClaimShunit Agmon, Amir Gilad, Brit Youngmann, Shahar Zoarets et al.VLDB 2025 · 5 citations
- Data-Semantics-Aware Recommendation of Diverse Pivot TablesWhanhee Cho, Anna FarihaSIGMOD 2026 · 4 citations
- Differentially Private Explanations for ClustersAmir Gilad, Tova Milo, Kathy Razmadze, Ron ZadicarioSIGMOD 2026 · 3 citations
- Fair and Actionable Causal Prescription RulesetBenton Li, Nativ Levy, Brit Youngmann, Sainyam Galhotra et al.SIGMOD 2025 · 3 citations
Builds on5
- Table2Analysis: Modeling and Recommendation of Common Analysis Patterns for Multi-Dimensional DataMengyu Zhou, Wang Tao, Pengxin Ji, Han Shi et al.AAAI 2020 · 26 citations
- Guided Exploration of User GroupsMariia Seleznova, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Eric SimonVLDB 2020 · 24 citations
- Summarizing Hierarchical Multidimensional DataAlexandra Kim, Laks V. S. Lakshmanan, Divesh SrivastavaICDE 2020 · 11 citations
- Exploring Ratings in Subjective DatabasesSihem Amer-Yahia, Tova Milo, Brit YoungmannSIGMOD 2021 · 8 citations
- Improving Constrained Search Results By Data MeliorationIdo Guy, Tova Milo, Slava Novgorodov, Brit YoungmannICDE 2021 · 2 citations
Related papers
- Reinforced Approximate Exploratory Data AnalysisShaddy Garg, Subrata Mitra, Tong Yu, Yash Gadhia et al.AAAI 2023 · 1 citation
- Supporting Guided Exploratory Visual Analysis on Time Series Data with Reinforcement LearningYang Shi, Bingchang Chen, Ying Chen, Zhuochen Jin et al.IEEE VIS 2023 · 8 citations
- Interpretable Attribute DiscretizationEugenie Lai, Inbal Croitoru, Brit Youngmann, Sainyam Galhotra et al.SIGMOD 2026
- Learning Opinion Summarizers by Selecting Informative ReviewsArthur Brazinskas, Mirella Lapata, Ivan TitovEMNLP 2021 · 26 citations
- Representative Query Results by VotingRachel Behar, Sara CohenSIGMOD 2022 · 2 citations
