Suna: Scalable Causal Confounder Discovery over Relational Data
Jiaxiang Liu, Siyuan Xia, Daniel Alabi, Eugene Wu
摘要
Understanding the causal relationships between treatments and outcomes is fundamental in various areas. Causal inference aims to estimate the effect of one variable on another, and critically relies on access to those variables as well as the key confounders. Unfortunately, data analysts often start with datasets lacking these columns, leading to incorrect estimations. Relational data repositories hold significant potential to augment such datasets with an admissible set of confounders necessary for causal analysis. While recent work has advocated for this potential, these approaches face notable limitations. They either assume the existence of a complete causal diagram over all datasets in the repository, which is impractical; rely on computationally infeasible techniques that do not scale to large data repositories with many features; or can only detect confounders in the absence of causal relations, and are thus ineffective when a causal effect exists. We observe that the asymmetry between causes and effects used in causal discovery can be exploited to directly identify confounders for causal queries. In this paper, we establish a connection between the existence of confounders and the presence of unconfounded ancestors of the treatment variable in the underlying causal diagram—without requiring access to the diagram. This makes it feasible to iteratively discover confounders until an admissible set is constructed. We propose Suna, a highly optimized, GPU-compatible system that implements a novel end-to-end algorithm for discovering confounders within large relational data repositories. Experiments on both real-world and synthetic datasets demonstrate that our system effectively discovers high-quality confounders. Furthermore, Suna employs algorithmic optimizations to accelerate confounder discovery without materializing joins. Our experiments show that Suna finds high-quality confounders while running >100x faster than existing confounder discovery systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
- Causal Relational LearningBabak Salimi, Harsh Parikh, Moe Kayali, Lise Getoor 等SIGMOD 2020 · 被引用 38 次
- A Sketch-based Index for Correlated Dataset SearchAécio S. R. Santos, Aline Bessa, Christopher Musco, Juliana FreireICDE 2022 · 被引用 31 次
- Metam: Goal-Oriented Data DiscoverySainyam Galhotra, Yue Gong, Raul Castro FernandezICDE 2023 · 被引用 28 次
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez 等VLDB 2020 · 被引用 14 次
相关 Paper
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 被引用 7 次
- Improving Visualization Interpretation Using CounterfactualsSmiti Kaul, David Borland, Nan Cao, David GotzIEEE VIS 2021 · 被引用 26 次
- Query-Specific Causal Graph Pruning Under Tiered KnowledgeYizuo Chen, Jane BarkerICLR 2026
- Detecting and Measuring Confounding Using Causal Mechanism ShiftsAbbavaram Gowtham Reddy, Vineeth N. BalasubramanianNeurIPS 2024 · 被引用 7 次
- Causal Data IntegrationBrit Youngmann, Michael J. Cafarella, Babak Salimi, Anna ZengVLDB 2023 · 被引用 14 次
