FEDEX: An Explainability Framework for Data Exploration Steps
Daniel Deutch, Amir Gilad, Tova Milo, Amit Mualem, Amit Somech
摘要
When exploring a new dataset, Data Scientists often apply analysis queries, look for insights in the resulting dataframe, and repeat to apply further queries. We propose in this paper a novel solution that assists data scientists in this laborious process. In a nutshell, our solution pinpoints the most interesting (sets of) rows in each obtained dataframe. Uniquely, our definition of interest is based on the contribution of each row to the interestingness of different columns of the entire dataframe, which, in turn, is defined using standard measures such as diversity and exceptionality. Intuitively, interesting rows are ones that explain why (some column of) the analysis query result is interesting as a whole. Rows are correlated in their contribution and so the interesting score for a set of rows may not be directly computed based on that of individual rows. We address the resulting computational challenge by restricting attention to semantically-related sets, based on multiple notions of semantic relatedness; these sets serve as more informative explanations. Our experimental study across multiple real-world datasets shows the usefulness of our system in various scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 被引用 12 次
- Finding Convincing Views to Endorse a ClaimShunit Agmon, Amir Gilad, Brit Youngmann, Shahar Zoarets 等VLDB 2025 · 被引用 5 次
- Differentially Private Explanations for ClustersAmir Gilad, Tova Milo, Kathy Razmadze, Ron ZadicarioSIGMOD 2026 · 被引用 3 次
- Fair and Actionable Causal Prescription RulesetBenton Li, Nativ Levy, Brit Youngmann, Sainyam Galhotra 等SIGMOD 2025 · 被引用 3 次
- Causal Explanations for Disparate Trends: Where and Why?Tal Blau, Brit Youngmann, Anna Fariha, Yuval MoskovitchSIGMOD 2026 · 被引用 2 次
它引用的顶会 Paper6
- Calliope: Automatic Visual Data Story Generation from a SpreadsheetDanqing Shi, Xinyue Xu, Fuling Sun, Yang Shi 等IEEE VIS 2020 · 被引用 179 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
- On Detecting Cherry-picked TrendlinesAbolfazl Asudeh, H. V. Jagadish, You Wu, Cong YuVLDB 2020 · 被引用 32 次
- Guided Exploration of User GroupsMariia Seleznova, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Eric SimonVLDB 2020 · 被引用 24 次
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic 等SIGMOD 2021 · 被引用 16 次
相关 Paper
- Selecting Sub-tables for Data ExplorationYael Amsterdamer, Susan B. Davidson, Tova Milo, Kathy Razmadze 等ICDE 2023 · 被引用 5 次
- Data-Semantics-Aware Recommendation of Diverse Pivot TablesWhanhee Cho, Anna FarihaSIGMOD 2026 · 被引用 4 次
- Representative Query Results by VotingRachel Behar, Sara CohenSIGMOD 2022 · 被引用 2 次
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 被引用 7 次
- Novel Table SearchBesat Kassaie, Renée J. MillerICDE 2026
