Suna: Scalable Causal Confounder Discovery over Relational Data
Jiaxiang Liu, Siyuan Xia, Daniel Alabi, Eugene Wu
Abstract
Understanding the causal relationships between treatments and outcomes is fundamental in various areas. Causal inference aims to estimate the effect of one variable on another, and critically relies on access to those variables as well as the key confounders. Unfortunately, data analysts often start with datasets lacking these columns, leading to incorrect estimations. Relational data repositories hold significant potential to augment such datasets with an admissible set of confounders necessary for causal analysis. While recent work has advocated for this potential, these approaches face notable limitations. They either assume the existence of a complete causal diagram over all datasets in the repository, which is impractical; rely on computationally infeasible techniques that do not scale to large data repositories with many features; or can only detect confounders in the absence of causal relations, and are thus ineffective when a causal effect exists. We observe that the asymmetry between causes and effects used in causal discovery can be exploited to directly identify confounders for causal queries. In this paper, we establish a connection between the existence of confounders and the presence of unconfounded ancestors of the treatment variable in the underlying causal diagram—without requiring access to the diagram. This makes it feasible to iteratively discover confounders until an admissible set is constructed. We propose Suna, a highly optimized, GPU-compatible system that implements a novel end-to-end algorithm for discovering confounders within large relational data repositories. Experiments on both real-world and synthetic datasets demonstrate that our system effectively discovers high-quality confounders. Furthermore, Suna employs algorithmic optimizations to accelerate confounder discovery without materializing joins. Our experiments show that Suna finds high-quality confounders while running >100x faster than existing confounder discovery systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 59 citations
- Causal Relational LearningBabak Salimi, Harsh Parikh, Moe Kayali, Lise Getoor et al.SIGMOD 2020 · 38 citations
- A Sketch-based Index for Correlated Dataset SearchAécio S. R. Santos, Aline Bessa, Christopher Musco, Juliana FreireICDE 2022 · 31 citations
- Metam: Goal-Oriented Data DiscoverySainyam Galhotra, Yue Gong, Raul Castro FernandezICDE 2023 · 28 citations
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez et al.VLDB 2020 · 14 citations
Related papers
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 7 citations
- Improving Visualization Interpretation Using CounterfactualsSmiti Kaul, David Borland, Nan Cao, David GotzIEEE VIS 2021 · 26 citations
- Query-Specific Causal Graph Pruning Under Tiered KnowledgeYizuo Chen, Jane BarkerICLR 2026
- Detecting and Measuring Confounding Using Causal Mechanism ShiftsAbbavaram Gowtham Reddy, Vineeth N. BalasubramanianNeurIPS 2024 · 7 citations
- Causal Data IntegrationBrit Youngmann, Michael J. Cafarella, Babak Salimi, Anna ZengVLDB 2023 · 14 citations
