Do Contemporary Causal Inference Models Capture Real-World Heterogeneity? Findings from a Large-Scale Benchmark
Haining Yu, Yizhou Sun
Abstract
We present unexpected findings from a large-scale benchmark study evaluating Conditional Average Treatment Effect (CATE) estimation algorithms, i.e., CATE models. By running 16 modern CATE models on 12 datasets and 43,200 sampled variants generated through diverse observational sampling strategies, we find that: (a) 62% of CATE estimates have a higher Mean Squared Error (MSE) than a trivial zero-effect predictor, rendering them ineffective; (b) in datasets with at least one useful CATE estimate, 80% still have higher MSE than a constant-effect model; and (c) Orthogonality-based models outperform other models only 30% of the time, despite widespread optimism about their performance. These findings highlight significant challenges in current CATE models and underscore the need for broader evaluation and methodological improvements. Our findings stem from a novel application of observational sampling, originally developed to evaluate Average Treatment Effect (ATE) estimates from observational methods with experiment data. To adapt observational sampling for CATE evaluation, we introduce a statistical parameter, Q, equal to MSE minus a constant and preserves the ranking of models by their MSE. We then derive a family of sample statistics, collectively called Q, that can be computed from real-world data. When used in observational sampling, Q is an unbiased estimator of Q and asymptotically selects the model with the smallest MSE. To ensure the benchmark reflects real-world heterogeneity, we handpick datasets where outcomes come from field rather than simulation. By integrating observational sampling, new statistics, and real-world datasets, the benchmark provides new insights into CATE model performance and reveals gaps in capturing real-world heterogeneity, emphasizing the need for more robust benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71d6d88c-cfd4-4a82-b812-7f3599bc46b8Builds on4
- Validating Causal Inference MethodsHarsh Parikh, Carlos Varjao, Louise Xu, Eric Tchetgen TchetgenICML 2022 · 36 citations
- In Search of Insights, Not Magic Bullets: Towards Demystification of the Model Selection Dilemma in Heterogeneous Treatment Effect EstimationAlicia Curth, Mihaela van der SchaarICML 2023 · 36 citations
- Empirical Analysis of Model Selection for Heterogeneous Causal Effect EstimationDivyat Mahajan, Ioannis Mitliagkas, Brady Neal, Vasilis SyrgkanisICLR 2024 · 28 citations
- How and Why to Use Experimental Data to Evaluate Methods for Observational Causal InferenceAmanda Gentzel, Purva Pruthi, David D. JensenICML 2021 · 22 citations
Related papers
- Comparison of meta-learners for estimating multi-valued treatment heterogeneous effectsNaoufal Acharki, Ramiro Lugo, Antoine Bertoncello, Josselin GarnierICML 2023 · 18 citations
- Counterfactual Cross-Validation: Stable Model Selection Procedure for Causal Inference ModelsYuta Saito, Shota YasuiICML 2020 · 34 citations
- A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured ConfoundingXinhai Zhang, Xingye QiaoNeurIPS 2024 · 1 citation
- Conditional Outcome Equivalence: A Quantile Alternative to CATEJosh Givens, Henry W. J. Reeve, Song Liu, Katarzyna RelugaNeurIPS 2024 · 3 citations
- Good Allocations from Bad EstimatesSílvia Casacuberta, Moritz HardtICLR 2026 · 3 citations
