Progressive Generalization Risk Reduction for Data-Efficient Causal Effect Estimation
Hechuan Wen, Tong Chen, Guanhua Ye, Li Kheng Chai, Shazia Sadiq, Hongzhi Yin
Abstract
Causal effect estimation (CEE) provides a crucial tool for predicting the unobserved counterfactual outcome for an entity. As CEE relaxes the requirement for "perfect" counterfactual samples (e.g., patients with identical attributes and only differ in treatments received) that are impractical to obtain and can instead operate on observational data, it is usually used in high-stake domains like medical treatment effect prediction. Nevertheless, in those high-stake domains, gathering a decently sized, fully labelled observational dataset remains challenging due to hurdles associated with costs, ethics, expertise and time needed, etc., of which medical treatment surveys are a typical example. Consequently, if the training dataset is small in scale, low generalization risks can hardly be achieved on any CEE algorithms. Unlike existing CEE methods that assume the constant availability of a dataset with abundant samples, in this paper, we study a more realistic CEE setting where the labelled data samples are scarce at the beginning, while more can be gradually acquired over the course of training -assuredly under a limited budget considering their expensive nature. Then, the problem naturally comes down to actively selecting the best possible samples to be labelled, e.g., identifying the next subset of patients to conduct the treatment survey. However, acquiring quality data for reducing the CEE risk under limited labelling budgets remains under-explored until now. To fill the gap, we theoretically analyse the generalization risk from an intriguing perspective of progressively shrinking its upper bound, and develop a principled label acquisition pipeline exclusively for CEE tasks. With our analysis, we propose the Model Agnostic Causal Active Learning (MACAL) algorithm for batchwise label acquisition, which aims to reduce both the CEE model's uncertainty and the post-acquisition distributional imbalance simultaneously at each acquisition step. Extensive experiments are
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ActiveCQ: Active Estimation of Causal QuantitiesErdun Gao, Dino SejdinovicICLR 2026 · 1 citation
- Treatment Effect Estimation with Differentiated Networked Effect on Graph DataXiaofeng Lin, Han Bao, Hisashi KashimaKDD 2026
- Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering PerspectiveHechuan Wen, Tong Chen, Mingming Gong, Li Kheng Chai et al.ICML 2025
Builds on8
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Gone Fishing: Neural Active Learning with Fisher EmbeddingsJordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, Sham M. KakadeNeurIPS 2021 · 124 citations
- Identifying Causal-Effect Inference Failure with Uncertainty-Aware ModelsAndrew Jesson, Sören Mindermann, Uri Shalit, Yarin GalNeurIPS 2020 · 85 citations
- Optimal Transport for Treatment Effect EstimationHao Wang, Jiajun Fan, Zhichao Chen, Haoxuan Li et al.NeurIPS 2023 · 71 citations
- Quantifying Ignorance in Individual-Level Causal-Effect Estimates under Hidden ConfoundingAndrew Jesson, Sören Mindermann, Yarin Gal, Uri ShalitICML 2021 · 66 citations
Related papers
- Budgeted Heterogeneous Treatment Effect EstimationTian Qin, Tian-Zuo Wang, Zhi-Hua ZhouICML 2021 · 18 citations
- Domain-wise Data Acquisition to Improve Performance under Distribution ShiftYue He, Dongbai Li, Pengfei Tian, Han Yu et al.ICML 2024 · 4 citations
- Active Learning with LLMs for Partially Observed and Cost-Aware ScenariosNicolás Astorga, Tennison Liu, Nabeel Seedat, Mihaela van der SchaarNeurIPS 2024 · 11 citations
- Active Policy Optimization for Individualized Dosing via Gradient Variance MinimizationYi Wan, Xin Wang, Huanhuan ChenICML 2026
- Batch Active Learning at ScaleGui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas et al.NeurIPS 2021 · 220 citations
