Observationally Informed Adaptive Causal Experimental Design
Erdun Gao, Liang Zhang, Jake Fawkes, Aoqi Zuo, Wenqin Liu, Haoxuan Li, Mingming Gong, Dino Sejdinovic
Abstract
Randomized Controlled Trials (RCTs) represent the gold standard for causal inference yet remain a scarce resource. While large-scale observational data is often available, it is typically used only for retrospective fusion, and remains discarded in prospective trial design due to bias concerns. We argue that this ''tabula rasa'' data acquisition strategy is inefficient when observational models contain useful structural information. In this work, we propose Active Residual Learning, a new paradigm that leverages the observational model as an informative but biased prior. This shifts the experimental focus from learning target causal quantities from scratch to estimating residual corrections that debias the observational model. To operationalize this, we introduce the R-Design framework. Theoretically, we characterize two key advantages: (1) a conditional structural efficiency gap, showing that estimating lower-complexity residual contrasts can admit faster convergence rates than reconstructing full outcomes; and (2) information efficiency, where we quantify the redundancy in standard parameter-based acquisition, demonstrating that such baselines can waste budget on task-irrelevant nuisance uncertainty. We propose R-EPIG (Residual Expected Predictive Information Gain), a unified criterion that directly targets the downstream causal quantity, reducing residual uncertainty for estimation or clarifying decision boundaries for policy. Experiments on synthetic and semi-synthetic benchmarks show that R-Design significantly outperforms baselines in the intended informative-but-biased regime, supporting the effectiveness of correcting a biased model rather than learning from scratch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 117683d3-547b-4ec1-9d2a-faaa0d024cdaBuilds on18
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 96 citations
- Active Bayesian Causal InferenceChristian Toth, Lars Lorch, Christian Knoll, Andreas Krause et al.NeurIPS 2022 · 52 citations
- CausalPFN: Amortized Causal Effect Estimation via In-Context LearningVahid Balazadeh Meresht, Hamidreza Kamkari, Valentin Thomas, Junwei Ma et al.NeurIPS 2025 · 52 citations
- Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational DataAndrew Jesson, Panagiotis Tigas, Joost van Amersfoort, Andreas Kirsch et al.NeurIPS 2021 · 42 citations
- Transductive Active Learning: Theory and ApplicationsJonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As et al.NeurIPS 2024 · 24 citations
Related papers
- Causal-EPIG: Causally Aligned Active CATE EstimationErdun Gao, Jake Fawkes, Dino SejdinovicICML 2026 · 3 citations
- ActiveCQ: Active Estimation of Causal QuantitiesErdun Gao, Dino SejdinovicICLR 2026 · 1 citation
- Budgeted Active Experimentation for Treatment Effect Estimation from Observational and Randomized DataJiacan Gao, Xinyan Su, Mingyuan Ma, Yiyan HUANG et al.ICML 2026 · 1 citation
- ABC3: Active Bayesian Causal Inference with Cohn Criteria in Randomized ExperimentsTaehun Cha, Donghun LeeAAAI 2025
- Learning-To-Measure: In-Context Active Feature AcquisitionYuta Kobayashi, Zilin Jing, Jiayu Yao, Hongseok Namkoong et al.ICML 2026 · 2 citations
