Out-of-Variable Generalisation for Discriminative Models
Siyuan Guo, Jonas Bernhard Wildberger, Bernhard Schölkopf
Abstract
The ability of an agent to do well in new environments is a critical aspect of intelligence. In machine learning, this ability is known as strong or out-of-distribution generalization. However, merely considering differences in distributions is inadequate for fully capturing differences between learning environments. In the present paper, we investigate out-of-variable generalization, which pertains to an agent's generalization capabilities concerning environments with variables that were never jointly observed before. This skill closely reflects the process of animate learning: we, too, explore Nature by probing, observing, and measuring proper subsets of variables at any given time. Mathematically, oov generalization requires the efficient re-use of past marginal information, i.e., information over subsets of previously observed variables. We study this problem, focusing on prediction tasks across environments that contain overlapping, yet distinct, sets of causes. We show that after fitting a classifier, the residual distribution in one environment reveals the partial derivative of the true generating function with respect to the unobserved causal parent in that environment. We leverage this information and propose a method that exhibits non-trivial out-of-variable generalization performance when facing an overlapping, yet distinct, set of causal predictors. Code:
Much of modern machine learning can be viewed as large-scale pattern recognition on suitably collected independent and identically distributed (i.i.d.) data. Its success builds on generalizing from one observation to the next, sampled from the same distribution. Animate intelligence differs from this in its ability to generalize from one problem to another. The machine learning community studies the latter under the term out-of-distribution (OOD) generalization
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label ExploitationPeng Tan, Hai-Tian Liu, Zhi-Hao Tan, Zhi-Hua ZhouNeurIPS 2024 · 8 citations
- Counterfactual reasoning: an analysis of in-context emergenceMoritz Miller, Bernhard Schölkopf, Siyuan GuoNeurIPS 2025 · 5 citations
- Sufficient Invariant Learning for Distribution ShiftTaero Kim, Subeen Park, Sungjun Lim, Yonghan Jung et al.CVPR 2025
- Learning Joint Interventional Effects from Single-Variable Interventions in Additive ModelsArmin Kekic, Sergio Hernan Garrido Mejia, Bernhard SchölkopfICML 2025
- Identifiable Exchangeable Mechanisms for Causal Structure and Representation LearningPatrik Reizinger, Siyuan Guo, Ferenc Huszár, Bernhard Schölkopf et al.ICLR 2025
Builds on13
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationKartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet et al.NeurIPS 2021 · 372 citations
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Learning explanations that are hard to varyGiambattista Parascandolo, Alexander Neitz, Antonio Orvieto, Luigi Gresele et al.ICLR 2021 · 221 citations
- Invariant Causal Representation Learning for Out-of-Distribution GeneralizationChaochao Lu, Yuhuai Wu, José Miguel Hernández-Lobato, Bernhard SchölkopfICLR 2022 · 119 citations
Related papers
- Out-of-distribution Generalization with Causal Invariant TransformationsRuoyu Wang, Mingyang Yi, Zhitang Chen, Shengyu ZhuCVPR 2022 · 40 citations
- Transportability for Bandits with Data from Different EnvironmentsAlexis Bellot, Alan Malek, Silvia ChiappaNeurIPS 2023 · 11 citations
- Improving Generalization of Dynamic Graph Learning via Environment PromptKuo Yang, Zhengyang Zhou, Qihe Huang, Limin Li et al.NeurIPS 2024 · 14 citations
- The Role of Pretrained Representations for the OOD Generalization of RL AgentsFrederik Träuble, Andrea Dittadi, Manuel Wuthrich, Felix Widmaier et al.ICLR 2022 · 19 citations
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau et al.ICML 2022 · 128 citations
