Putting Things into Context: Rich Explanations for Query Answers using Join Graphs
Chenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic, Sudeepa Roy
Abstract
In many data analysis applications, there is a need to explain why a surprising or interesting result was produced by a query. Previous approaches to explaining results have directly or indirectly used data provenance (input tuples contributing to the result(s) of interest), which is limited by the fact that relevant information for explaining an answer may not be fully contained in the provenance. We propose a new approach for explaining query results by augmenting provenance with information from other related tables in the database. We develop a suite of optimization techniques, and demonstrate experimentally using real datasets and through a user study that our approach produces meaningful results by efficiently navigating the large search space of possible explanations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0fc6c25-4b63-47bb-8787-0f6441afb3b1Cited by top-tier papers12
- Why Not Yet: Fixing a Top-k Ranking that Is Not Fair to IndividualsZixuan Chen, Panagiotis Manolios, Mirek RiedewaldVLDB 2023 · 17 citations
- DPXPlain: Privately Explaining Aggregate Query AnswersYuchao Tao, Amir Gilad, Ashwin Machanavajjhala, Sudeepa RoyVLDB 2023 · 15 citations
- FEDEX: An Explainability Framework for Data Exploration StepsDaniel Deutch, Amir Gilad, Tova Milo, Amit Mualem et al.VLDB 2022 · 15 citations
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 12 citations
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 7 citations
Builds on3
- Approximate Summaries for Why and Why-not ProvenanceSeokki Lee, Bertram Ludäscher, Boris GlavicVLDB 2020 · 29 citations
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez et al.VLDB 2020 · 14 citations
- Summarizing Hierarchical Multidimensional DataAlexandra Kim, Laks V. S. Lakshmanan, Divesh SrivastavaICDE 2020 · 11 citations
Related papers
- On Optimizing the Trade-off between Privacy and Utility in Data ProvenanceDaniel Deutch, Ariel Frankenthal, Amir Gilad, Yuval MoskovitchSIGMOD 2021 · 15 citations
- Computing the Why-Provenance for Datalog Queries via SAT SolversMarco Calautti, Ester Livshits, Andreas Pieris, Markus SchneiderAAAI 2024 · 4 citations
- To Not Miss the Forest for the Trees - A Holistic Approach for Explaining Missing Answers over Nested DataRalf Diestelkämper, Seokki Lee, Melanie Herschel, Boris GlavicSIGMOD 2021 · 15 citations
- Explaining Inference Queries with Bayesian OptimizationBrandon Lockhart, Jinglin Peng, Weiyuan Wu, Jiannan Wang et al.VLDB 2021 · 9 citations
- Computing the Shapley Value of Facts in Query AnsweringDaniel Deutch, Nave Frost, Benny Kimelfeld, Mikaël MonetSIGMOD 2022 · 31 citations
