Learning Adjustment Sets from Observational and Limited Experimental Data
Sofia Triantafillou, Gregory F. Cooper
Abstract
We can estimate causal effects from observational data if an appropriate set of covariates (an adjustment set) can be identified, which removes confounding bias; however, such a set is often not identifiable from observational data alone. Experimental data allow unbiased causal effect estimation, but are typically limited in sample size and can therefore yield estimates of high variance. Moreover, experiments are often performed on a different (specialized) population than the population of interest. In this work, we introduce a method that combines large observational and limited experimental data to identify adjustment sets and improve the estimation of causal effects for a target population. The method scores an adjustment set by calculating the marginal likelihood for the experimental data given an observationally-derived causal effect estimate, using a putative adjustment set. The method can make inferences that are not possible using constraintbased methods. We show that the method can improve causal effect estimation, and can make additional inferences when compared to state-of-the-art methods. In this work, we present a method for using predictions of these ID estimates (from observational data) 46 as priors in deriving the marginal likelihood of the experimental data in order to score adjustment sets 47 and find the most probable one. By using the observational and experimental data in this way, our 48 method can make inferences that are not possible based on conditional (in) dependence constraints 49 alone, like identifying that C is an adjustment set for D, AE in Fig. 1 , even though C is not measured 50 in the experimental data. To our knowledge, this method is the first one described in the literature 51 that can make this inference. 2 Preliminaries 53 We use the framework of Semi-Markov Causal Models (SMCMs). We assume the reader is familiar 54 with of causal graphical models and related terminology. We use the terms node and variable 55 interchangeably. We use bold to denote variable sets, uppercase letters to denote single variables, 56 and lowercase letters to denote variable values. If we know the causal SMCM G, a hard intervention 57 of where a treatment X is set to x can be represented with the do-operator, do(X=x). The post 58 intervention distribution of an outcome Y given do(X=x) is denoted P (Y |do(X=x) or P X (Y ). In 59
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9c0b32f-6b16-440c-a705-e9d5e5d7fbcfCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Estimating Possible Causal Effects with Latent Variables via AdjustmentTian-Zuo Wang, Tian Qin, Zhi-Hua ZhouICML 2023 · 16 citations
- Fast Proxy Experiment Design for Causal Effect IdentificationSepehr Elahi, Sina Akbari, Jalal Etesami, Negar Kiyavash et al.NeurIPS 2024 · 3 citations
- Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent VariablesZheng Li, Xichen Guo, Feng Xie, Yan Zeng et al.NeurIPS 2025 · 4 citations
- Bounds on Causal Effects and Application to High Dimensional DataAng Li, Judea PearlAAAI 2022 · 25 citations
- Optimal Adjustment Sets for Nonparametric Estimation of Weighted Controlled Direct EffectRuiyang Lin, Yongyi Guo, Kyra GanNeurIPS 2025 · 3 citations
