Invariant Causal Representation Learning for Out-of-Distribution Generalization
Chaochao Lu, Yuhuai Wu, José Miguel Hernández-Lobato, Bernhard Schölkopf
Abstract
Due to spurious correlations, machine learning systems often fail to generalize to environments whose distributions differ from the ones used at training time. Prior work addressing this, either explicitly or implicitly, attempted to find a data representation that has an invariant relationship with the target. This is done by leveraging a diverse set of training environments to reduce the effect of spurious features and build an invariant predictor. However, these methods have generalization guarantees only when both data representation and classifiers come from a linear model class. We propose invariant Causal Representation Learning (iCaRL), an approach that enables out-of-distribution (OOD) generalization in the nonlinear setting (i.e., nonlinear representations and nonlinear classifiers). It builds upon a practical and general assumption: the prior over the data representation (i.e., a set of latent variables encoding the data) given the target and the environment belongs to general exponential family distributions, i.e., a more flexible conditionally non-factorized prior that can actually capture complicated dependences between the latent variables. Based on this, we show that it is possible to identify the data representation up to simple transformations. We also show that all direct causes of the target can be fully discovered, which further enables us to obtain generalization guarantees in the nonlinear setting. Experiments on both synthetic and real-world datasets demonstrate that our approach outperforms a variety of baseline methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers41
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele et al.NeurIPS 2023 · 127 citations
- Unleashing the Power of Graph Data Augmentation on Covariate Distribution ShiftYongduo Sui, Qitian Wu, Jiancan Wu, Qing Cui et al.NeurIPS 2023 · 63 citations
- Joint Learning of Label and Environment Causal Independence for Graph Out-of-Distribution GeneralizationShurui Gui, Meng Liu, Xiner Li, Youzhi Luo et al.NeurIPS 2023 · 54 citations
- Leveraging sparse and shared feature activations for disentangled representation learningMarco Fumero, Florian Wenzel, Luca Zancato, Alessandro Achille et al.NeurIPS 2023 · 42 citations
- Improving Non-Transferable Representation Learning by Harnessing Content and StyleZiming Hong, Zhenyi Wang, Li Shen, Yu Yao et al.ICLR 2024 · 37 citations
Related papers
- Identifying Representations for Intervention ExtrapolationSorawit Saengkyongam, Elan Rosenfeld, Pradeep Kumar Ravikumar, Niklas Pfister et al.ICLR 2024 · 20 citations
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
- Diagnosing and Rectifying Fake OOD Invariance: A Restructured Causal ApproachZiliang Chen, Yongsen Zheng, Zhao-Rong Lai, Quanlong Guan et al.AAAI 2024 · 4 citations
- Out-of-distribution Generalization with Causal Invariant TransformationsRuoyu Wang, Mingyang Yi, Zhitang Chen, Shengyu ZhuCVPR 2022 · 40 citations
- Causal Transportability for Visual RecognitionChengzhi Mao, Kevin Xia, James Wang, Hao Wang et al.CVPR 2022 · 27 citations
