FEDEX: An Explainability Framework for Data Exploration Steps
Daniel Deutch, Amir Gilad, Tova Milo, Amit Mualem, Amit Somech
Abstract
When exploring a new dataset, Data Scientists often apply analysis queries, look for insights in the resulting dataframe, and repeat to apply further queries. We propose in this paper a novel solution that assists data scientists in this laborious process. In a nutshell, our solution pinpoints the most interesting (sets of) rows in each obtained dataframe. Uniquely, our definition of interest is based on the contribution of each row to the interestingness of different columns of the entire dataframe, which, in turn, is defined using standard measures such as diversity and exceptionality. Intuitively, interesting rows are ones that explain why (some column of) the analysis query result is interesting as a whole. Rows are correlated in their contribution and so the interesting score for a set of rows may not be directly computed based on that of individual rows. We address the resulting computational challenge by restricting attention to semantically-related sets, based on multiple notions of semantic relatedness; these sets serve as more informative explanations. Our experimental study across multiple real-world datasets shows the usefulness of our system in various scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d89cf1d-9b28-4fbc-9e89-2bcb6e5f98a5Cited by top-tier papers6
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 12 citations
- Finding Convincing Views to Endorse a ClaimShunit Agmon, Amir Gilad, Brit Youngmann, Shahar Zoarets et al.VLDB 2025 · 5 citations
- Differentially Private Explanations for ClustersAmir Gilad, Tova Milo, Kathy Razmadze, Ron ZadicarioSIGMOD 2026 · 3 citations
- Fair and Actionable Causal Prescription RulesetBenton Li, Nativ Levy, Brit Youngmann, Sainyam Galhotra et al.SIGMOD 2025 · 3 citations
- Causal Explanations for Disparate Trends: Where and Why?Tal Blau, Brit Youngmann, Anna Fariha, Yuval MoskovitchSIGMOD 2026 · 2 citations
Builds on6
- Calliope: Automatic Visual Data Story Generation from a SpreadsheetDanqing Shi, Xinyue Xu, Fuling Sun, Yang Shi et al.IEEE VIS 2020 · 179 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
- On Detecting Cherry-picked TrendlinesAbolfazl Asudeh, H. V. Jagadish, You Wu, Cong YuVLDB 2020 · 32 citations
- Guided Exploration of User GroupsMariia Seleznova, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Eric SimonVLDB 2020 · 24 citations
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic et al.SIGMOD 2021 · 16 citations
Related papers
- Selecting Sub-tables for Data ExplorationYael Amsterdamer, Susan B. Davidson, Tova Milo, Kathy Razmadze et al.ICDE 2023 · 5 citations
- Data-Semantics-Aware Recommendation of Diverse Pivot TablesWhanhee Cho, Anna FarihaSIGMOD 2026 · 4 citations
- Representative Query Results by VotingRachel Behar, Sara CohenSIGMOD 2022 · 2 citations
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 7 citations
- Novel Table SearchBesat Kassaie, Renée J. MillerICDE 2026
