Causal Explanations for Disparate Trends: Where and Why?
Tal Blau, Brit Youngmann, Anna Fariha, Yuval Moskovitch
Abstract
During data analysis, we are often perplexed by certain disparities observed between two groups of interest within a dataset. To better understand an observed disparity, we need explanations that can pinpoint the data regions where the disparity is most pronounced, along with its causes, i.e., factors that alleviate or exacerbate the disparity. This task can be complex and tedious, particularly when the dataset is large and high-dimensional, demanding an automatic system for discovering explanations (data regions and causes) of an observed disparity in a dataset. When offering explanations for disparities, it is critical that they are not only interpretable but also actionable—enabling users to make informed, data-driven decisions. This requires explanations to go beyond surface-level correlations and instead capture causal relationships. We introduce ExDis, a framework for discovering causal Ex planations for Dis parities between two groups of interest. ExDis identifies data regions (subpopulations) where disparities are most pronounced (or reversed), and associates specific factors that causally contribute to the disparity within each identified data region. We formally define the ExDis framework and the associated optimization problem, analyze its complexity, and develop an efficient algorithm to solve the problem. Through extensive experiments over three real-world datasets, we demonstrate that ExDis generates meaningful causal explanations, outperforms prior methods, and scales effectively to handle large, high-dimensional datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6809b287-81fb-4d0b-a0dd-d1ec34361348Builds on19
- Explaining Black-Box Algorithms Using Probabilistic Contrastive CounterfactualsSainyam Galhotra, Romila Pradhan, Babak SalimiSIGMOD 2021 · 85 citations
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 51 citations
- Approximate Summaries for Why and Why-not ProvenanceSeokki Lee, Bertram Ludäscher, Boris GlavicVLDB 2020 · 29 citations
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 22 citations
- XInsight: eXplainable Data Analysis Through The Lens of CausalityPingchuan Ma, Rui Ding, Shuai Wang, Shi Han et al.SIGMOD 2023 · 20 citations
Related papers
- Explaining Algorithmic Fairness Through Fairness-Aware Causal Path DecompositionWeishen Pan, Sen Cui, Jiang Bian, Changshui Zhang et al.KDD 2021 · 29 citations
- Feature Importance Disparities for Data Bias InvestigationsPeter W. Chang, Leor Fishman, Seth NeelICML 2024 · 3 citations
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 53 citations
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 12 citations
- Query-Driven Data Exploration with Heterogeneous Treatment EffectsAntonis Mandamadiotis, Sihem Amer-Yahia, Georgia KoutrikaICDE 2026
