When Shift Happens - Confounding Is to Blame
Abbavaram Gowtham Reddy, Celia Rubio-Madrigal, Rebekka Burkholz, Krikamol Muandet
Abstract
Distribution shifts introduce uncertainty that undermines the robustness and generalization capabilities of machine learning models. While conventional wisdom suggests that learning causal-invariant representations enhances robustness to such shifts, recent empirical studies present a counterintuitive finding: (i) empirical risk minimization (ERM) can rival or even outperform state-of-the-art out-ofdistribution (OOD) generalization methods, and (ii) its OOD generalization performance improves when all available covariates-not just causal ones-are utilized. Drawing on both empirical and theoretical evidence, we attribute this phenomenon to hidden confounding. Shifts in hidden confounding induce changes in data distributions that violate assumptions commonly made by existing OOD generalization approaches. Under such conditions, we prove that effective generalization requires learning environment-specific relationships, rather than relying solely on invariant ones. Furthermore, we show that models augmented with proxies for hidden confounders can mitigate the challenges posed by hidden confounding shifts. These findings offer new theoretical insights and practical guidance for designing robust OOD generalization algorithms and principled covariate selection strategies. All variable models vs causal models. Recently, Nastl and Hardt [57] introduced a benchmark study where covariates are categorized into four groups: causal (conservatively chosen), arguably causal, anti-causal, and other spurious covariates. They show that across 16 benchmark datasets, models using all covariates Pareto-dominate those using only causal or arguably causal subsets on both ID and OOD data. However, there is limited theoretical work explaining these results. We present scenarios and arguments to explain their experimental findings. For linear causal models, anchor regression [68] introduces a framework that balances between two estimation paradigms: models that include all observed covariates and models that focus solely on causal covariates. We aim to explain the impact of adding more covariates that are not necessarily causal under a hidden confounding shift. Eastwood et al. [21] show that unstable covariates can boost performance when they carry information about the label, provided they are conditionally independent of the stable covariates given the label. They propose to adjust the distribution shift by looking at the test domain without labels. However, when applied to a medical real-world dataset not constructed for this particular problem [8], ERM still remains competitive with their method, in line with the findings of Nastl and Hardt [57]. This reflects the broader insight that, under well-specified covariate shifts, maximum likelihood estimation (MLE) achieves minimax optimality for OOD generalization [26] . Yet, real-world settings are rarely well-specified due to hidden confounding shift, which is the main focus of this work. Manifestations of hidden confounding shift In this section, we provide background on hidden confounding shift and motivate the need to address it using an example involving causal effect identification of observed covariates on the outcome.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1b7b28e-61d9-4581-99a4-de8c82fa9eceCited by top-tier papers4
- An Analysis of Causal Effect Estimation using Outcome Invariant Data AugmentationUzair Akbar, Niki Kilbertus, Hao Shen, Krikamol Muandet et al.NeurIPS 2025 · 3 citations
- Fixed Aggregation Features Can Rival GNNsCelia Rubio-Madrigal, Rebekka BurkholzICML 2026 · 2 citations
- Counterfactual Residual Data Augmentation for RegressionHossein Mohebbi, Oliver Schulte, Ke Li, Pascal PoupartICML 2026
- Boosting for Predictive SufficiencyAbbavaram Gowtham Reddy, Rajeev Verma, Celia Rubio-Madrigal, Krikamol Muandet et al.ICLR 2026
Builds on26
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 755 citations
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann et al.NeurIPS 2020 · 688 citations
Related papers
- Empirical or Invariant Risk Minimization? A Sample Complexity PerspectiveKartik Ahuja, Jun Wang, Amit Dhurandhar, Karthikeyan Shanmugam et al.ICLR 2021 · 15 citations
- Distribution Shift Is Key to Learning Invariant PredictionHong Zheng, Fei TengAAAI 2026
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationKartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet et al.NeurIPS 2021 · 372 citations
- Generative multitask learning mitigates target-causing confoundingTaro Makino, Krzysztof J. Geras, Kyunghyun ChoNeurIPS 2022 · 9 citations
- Heterogeneous Risk MinimizationJiashuo Liu, Zheyuan Hu, Peng Cui, Bo Li et al.ICML 2021 · 170 citations
