Uncovering Bias Mechanisms in Observational Studies
Ilker Demirel, Zeshan Hussain, Piersilvio De Bartolomeis, David Sontag
Abstract
Observational studies are a key resource for causal inference but are often affected by systematic biases. Prior work has focused mainly on detecting these biases, via sensitivity analyses and comparisons with randomized controlled trials, or mitigating them through debiasing techniques. However, there remains a lack of methodology for uncovering the underlying mechanisms driving these biases, e.g., whether due to hidden confounding or selection of participants. In this work, we show that the relationship between bias magnitude and the predictive performance of nuisance function estimators (in the observational study) can help distinguish among common sources of causal bias. We validate our methodology through extensive synthetic experiments and a real-world case study, demonstrating its effectiveness in revealing the mechanisms behind observed biases. Our framework offers a new lens for understanding and characterizing bias in observational studies, with practical implications for improving causal inference. * Equal contribution. 2 We include an extended related work section in Appendix C. Preprint. Under review. recovering a single graph but instead groups together several graphs that exhibit equivalent bias patterns. To that end, we begin by estimating the bias function in the OS using data from a RCT conducted on a population supported in both datasets. We then examine how the bias varies across patients as a function of the performance of predictive models 3 fitted on the OS data. This relationship gives rise to empirical statistics that can effectively differentiate between distinct bias mechanisms. Contributions First, we establish a comprehensive taxonomy of common causal biases, each corresponding to a distinct class of graphs (Section 3). We describe a generative model for the OS in the context of these graphs motivated by clinical decision-making (Section 4.1). Next, we demonstrate a relationship between the predictive performance of the nuisance functions and the bias function in the OS (Section 4.2). To quantify this relationship, we use the covariance between the prediction error and magnitude of the bias, which enables provable discrimination of different bias mechanisms (Section 4.3). For practical implementation, we propose consistent estimators of the covariance (Section 4.5). Finally, we validate our methodology through both synthetic experiments and real-world analysis using data from Women's Health Initiative [59] (Sections 5 and 6). Notation and Background Let A denote a treatment action, X the set of measured patient covariates at baseline, and Y the observed outcome of interest. We denote by Y a the potential outcome under A = a. For each patient i, we observe only one of their potential outcomes, that is, We assume access to patient-level data from an RCT and an OS, and use R = 1 and R = 0 to represent the underlying RCT and OS populations, respectively. We use S to denote whether a patient was selected into the study cohort for analysis. For instance, a patient may be excluded from the analysis in an RCT if they did not adhere to their treatment assignment (i.e., R i = 1 and S i = 0). From a large insurance claims dataset, a patient maybe selected into the OS cohort when emulating an RCT if they meet the eligibility criteria (i.e., R i = 0 and S i = 1). We assume that X is available for all patients, and that A and Y are available for those selected into the analysis (S i = 1). Finally, we let U denote the set of unmeasured covariates that can influence the downstream variables S, A, and Y a in the OS. Such omitted variables are the reasons behind many common causal biases in observational studies, which are described in Section 3. The causal estimand we focus on is the conditional average treatment effect (CATE), defined as Estimating the CATE is challenging when unobserved covariates, U , influence treatment assignment (A), outcome (Y ), and selection into the study cohort (S). We list below necessary conditions to identify the CATE in a population, which are often satisfied in RCTs, but can be violated in OSes. Assumption 2.1 (Internal validity of RCT). The following hold in the RCT (R = 1) for all a ∈ 0, 1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcb2c7c2-4c4b-4bbe-ad6b-073c98e8d7e8Builds on10
- Recursive Causal Structure Learning in the Presence of Latent Variables and Selection BiasSina Akbari, Ehsan Mokhtarian, AmirEmad Ghassami, Negar KiyavashNeurIPS 2021 · 37 citations
- Detecting hidden confounding in observational data using multiple environmentsRickard Karlsson, Jesse H. KrijtheNeurIPS 2023 · 22 citations
- Verification and search algorithms for causal DAGsDavin Choo, Kirankumar Shiragur, Arnab BhattacharyyaNeurIPS 2022 · 21 citations
- Score-Based Causal Discovery of Latent Variable Causal ModelsIgnavier Ng, Xinshuai Dong, Haoyue Dai, Biwei Huang et al.ICML 2024 · 18 citations
- Prediction-powered Generalization of Causal InferencesIlker Demirel, Ahmed M. Alaa, Anthony Philippakis, David A. SontagICML 2024 · 18 citations
Related papers
- Learning Disentangled Representations for CounterFactual RegressionNegar Hassanpour, Russell GreinerICLR 2020 · 176 citations
- Sense and Sensitivity Analysis: Simple Post-Hoc Analysis of Bias Due to Unobserved ConfoundingVictor Veitch, Anisha ZaveriNeurIPS 2020 · 67 citations
- Towards a Holistic Understanding of Selection Bias for Causal Effect IdentificationYiwen (Evie) Qiu, Filip Kovačević, Shimeng Huang, Peter Spirtes et al.ICML 2026
- A Practical Upper Bound on Selection Bias Effects in Medical Prediction ModelsKara Liu, Maggie Wang, Russ B. AltmanKDD 2026
- When Selection Meets Intervention: Additional Complexities in Causal DiscoveryHaoyue Dai, Ignavier Ng, Jianle Sun, Zeyu Tang et al.ICLR 2025
