On Explaining Confounding Bias
Brit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak Salimi
Abstract
When analyzing large datasets, analysts are often interested in the explanations for surprising or unexpected results produced by their queries. In this work, we focus on aggregate SQL queries that expose correlations in the data. A major challenge that hinders the interpretation of such queries is confounding bias, which can lead to an unexpected correlation. We generate explanations in terms of a set of confounding variables that explain the unexpected correlation observed in a query. We propose to mine candidate confounding variables from external sources since, in many real-life scenarios, the explanations are not solely contained in the input data. We present an efficient algorithm that finds the optimal subset of attributes (mined from external sources and the input dataset) that explain the unexpected correlation. This algorithm is embodied in a system called MESA. We demonstrate experimentally over multiple real-life datasets and through a user study that our approach generates insightful explanations, outperforming existing methods that search for explanations only in the input data. We further demonstrate the robustness of our system to missing data and the ability of MESA to handle input datasets containing millions of tuples and an extensive search space of candidate confounding attributes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Causal Data IntegrationBrit Youngmann, Michael J. Cafarella, Babak Salimi, Anna ZengVLDB 2023 · 14 citations
- A Unified Approach for Resilience and Causal Responsibility with Integer Linear Programming (ILP) and LP RelaxationsNeha Makhija, Wolfgang GatterbauerSIGMOD 2024 · 13 citations
- Uncovering the Propensity Identification Problem in Debiased RecommendationsHonglei Zhang, Shuyi Wang, Haoxuan Li, Chunyuan Zheng et al.ICDE 2024 · 12 citations
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 12 citations
- Finding Convincing Views to Endorse a ClaimShunit Agmon, Amir Gilad, Brit Youngmann, Shahar Zoarets et al.VLDB 2025 · 5 citations
Builds on10
- Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental StudyFarahnaz Akrami, Mohammed Samiul Saeef, Qingheng Zhang, Wei Hu et al.SIGMOD 2020 · 101 citations
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 98 citations
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 53 citations
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco et al.SIGMOD 2021 · 45 citations
- Approximate Summaries for Why and Why-not ProvenanceSeokki Lee, Bertram Ludäscher, Boris GlavicVLDB 2020 · 29 citations
Related papers
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic et al.SIGMOD 2021 · 16 citations
- Suna: Scalable Causal Confounder Discovery over Relational DataJiaxiang Liu, Siyuan Xia, Daniel Alabi, Eugene WuVLDB 2025
- To Not Miss the Forest for the Trees - A Holistic Approach for Explaining Missing Answers over Nested DataRalf Diestelkämper, Seokki Lee, Melanie Herschel, Boris GlavicSIGMOD 2021 · 15 citations
- Explaining Inference Queries with Bayesian OptimizationBrandon Lockhart, Jinglin Peng, Weiyuan Wu, Jiannan Wang et al.VLDB 2021 · 9 citations
- "What makes my queries slow?": Subgroup Discovery for SQL Workload AnalysisYoucef Remil, Anes Bendimerad, Romain Mathonat, Philippe Chaleat et al.ASE 2021 · 12 citations
