A Theory of Independent Mechanisms for Extrapolation in Generative Models
Michel Besserve, Rémy Sun, Dominik Janzing, Bernhard Schölkopf
Abstract
Generative models can be trained to emulate complex empirical data, but are they useful to make predictions in the context of previously unobserved environments? An intuitive idea to promote such extrapolation capabilities is to have the architecture of such model reflect a causal graph of the true data generating process, such that one can intervene on each node independently of the others. However, the nodes of this graph are usually unobserved, leading to overparameterization and lack of identifiability of the causal structure. We develop a theoretical framework to address this challenging situation by defining a weaker form of identifiability, based on the principle of independence of mechanisms. We demonstrate on toy examples that classical stochastic gradient descent can hinder the model's extrapolation capabilities, suggesting independence of mechanisms should be enforced explicitly during training. Experiments on deep generative models trained on real world data support these insights and illustrate how the extrapolation capabilities of such models can be leveraged.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4abcb3a1-956a-4f0b-b774-5d227fc6b720Cited by top-tier papers10
- Visual Representation Learning Does Not Generalize Strongly Within the Same DomainLukas Schott, Julius von Kügelgen, Frederik Träuble, Peter Vincent Gehler et al.ICLR 2022 · 79 citations
- Additive Decoders for Latent Variables Identification and Cartesian-Product ExtrapolationSébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, Simon Lacoste-JulienNeurIPS 2023 · 61 citations
- Provably Learning Object-Centric RepresentationsJack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf et al.ICML 2023 · 55 citations
- Dynamic Inference with Neural InterpretersNasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter V. Gehler et al.NeurIPS 2021 · 35 citations
- Demystifying Inductive Biases for (Beta-)VAE Based ArchitecturesDominik Zietlow, Michal Rolínek, Georg MartiusICML 2021 · 24 citations
Builds on4
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- Gradient-Based Neural DAG LearningSébastien Lachapelle, Philippe Brouillard, Tristan Deleu, Simon Lacoste-JulienICLR 2020 · 337 citations
- Causal Discovery with Reinforcement LearningShengyu Zhu, Ignavier Ng, Zhitang ChenICLR 2020 · 285 citations
- Counterfactuals uncover the modular structure of deep generative modelsMichel Besserve, Arash Mehrjou, Rémy Sun, Bernhard SchölkopfICLR 2020 · 109 citations
Related papers
- Independent mechanism analysis, a new concept?Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf et al.NeurIPS 2021 · 133 citations
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 56 citations
- Identifiable Exchangeable Mechanisms for Causal Structure and Representation LearningPatrik Reizinger, Siyuan Guo, Ferenc Huszár, Bernhard Schölkopf et al.ICLR 2025
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele et al.NeurIPS 2023 · 127 citations
- Intervention Generalization: A View from Factor Graph ModelsGecia Bravo Hermsdorff, David S. Watson, Jialin Yu, Jakob Zeitler et al.NeurIPS 2023 · 7 citations
