Causal Balancing for Domain Generalization
Xinyi Wang, Michael Saxon, Jiachen Li, Hongyang Zhang, Kun Zhang, William Yang Wang
Abstract
While machine learning models rapidly advance the state-of-the-art on various real-world tasks, out-of-domain (OOD) generalization remains a challenging problem given the vulnerability of these models to spurious correlations. We propose a balanced mini-batch sampling strategy to transform a biased data distribution into a spurious-free balanced distribution, based on the invariance of the underlying causal mechanisms for the data generation process. We argue that the Bayes optimal classifiers trained on such balanced distribution are minimax optimal across a diverse enough environment space. We also provide an identifiability guarantee of the latent variable model of the proposed data generation process, when utilizing enough train environments. Experiments are conducted on DomainBed, demonstrating empirically that our method obtains the best performance across 20 baselines reported on the benchmark. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Flatness-Aware Minimization for Domain GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Yancheng Dong et al.ICCV 2023 · 37 citations
- Rethinking Misalignment in Vision-Language Model Adaptation from a Causal PerspectiveYanan Zhang, Jiangmeng Li, Lixiang Liu, Wenwen QiangNeurIPS 2024 · 16 citations
- Gradient Extrapolation for Debiased Representation LearningIhab Asaad, Maha Shadaydeh, Joachim DenzlerICCV 2025 · 4 citations
- CauDiTS: Causal Disentangled Domain Adaptation of Multivariate Time SeriesJunxin Lu, Shiliang SunICML 2024 · 2 citations
- Subgroups Matter for Robust Bias MitigationAnissa Alloula, Charles Jones, Ben Glocker, Bartlomiej W. PapiezICML 2025
Builds on26
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li et al.AAAI 2020 · 499 citations
Related papers
- On the Out-of-Distribution Generalization of Self-Supervised LearningWenwen Qiang, Jingyao Wang, Zeen Song, Jiangmeng Li et al.ICML 2025
- Out-of-distribution Generalization with Causal Invariant TransformationsRuoyu Wang, Mingyang Yi, Zhitang Chen, Shengyu ZhuCVPR 2022 · 40 citations
- Lost Domain Generalization Is a Natural Consequence of Lack of Training DomainsYimu Wang, Yihan Wu, Hongyang ZhangAAAI 2024 · 7 citations
- Breaking Correlation Shift via Conditional Invariant RegularizerMingyang Yi, Ruoyu Wang, Jiacheng Sun, Zhenguo Li et al.ICLR 2023
- On Calibration and Out-of-Domain GeneralizationYoav Wald, Amir Feder, Daniel Greenfeld, Uri ShalitNeurIPS 2021 · 184 citations
