Learning Linear Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity
Jikai Jin, Vasilis Syrgkanis
Abstract
We study causal representation learning, the task of recovering high-level latent variables and their causal relationships in the form of a causal graph from low-level observed data (such as text and images), assuming access to observations generated from multiple environments. Prior results on the identifiability of causal representations typically assume access to single-node interventions which is rather unrealistic in practice, since the latent variables are unknown in the first place. In this work, we consider the task of learning causal representation learning with data collected from general environments . We show that even when the causal model and the mixing function are both linear, there exists a surrounded-node ambiguity (SNA) [46] which is basically unavoidable in our setting. On the other hand, in the same linear case, we show that identification up to SNA is possible under mild conditions, and propose an algorithm, LiNGCReL which provably achieves such identifiability guarantee. We conduct extensive experiments on synthetic data and demonstrate the effectiveness of LiNGCReL in the finite-sample regime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa8b0f92-4287-47d3-9ac2-3376f1f96deeCited by top-tier papers3
- World Models in Pieces: Structural Certification for General AgentsYikai Lu, Yifei Wu, Xinyu Lu, Tongxin LiICML 2026
- Identifying dependent components from multi-domain linear mixturesDanru Xu, Lauri Parkkonen, Sara Magliacane, Aapo HyvarinenICML 2026
- Reward-oriented Causal Representation LearningZirui Yan, Emre Acartürk, Ali TajerNeurIPS 2025
Builds on16
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf et al.ICML 2020 · 361 citations
- Weakly supervised causal representation learningJohann Brehmer, Pim de Haan, Phillip Lippe, Taco S. CohenNeurIPS 2022 · 196 citations
- Interventional Causal Representation LearningKartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua BengioICML 2023 · 143 citations
Related papers
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele et al.NeurIPS 2023 · 127 citations
- Sample Complexity of Interventional Causal Representation LearningEmre Acartürk, Burak Varici, Karthikeyan Shanmugam, Ali TajerNeurIPS 2024 · 9 citations
- Linear Causal Representation Learning from Unknown Multi-node InterventionsBurak Varici, Emre Acartürk, Karthikeyan Shanmugam, Ali TajerNeurIPS 2024 · 19 citations
- Linear Causal Representation Learning by Topological Ordering, Pruning, and DisentanglementHao Chen, Lin Liu, Yuguang WangICML 2026
- Identifiable Latent Polynomial Causal Models through the Lens of ChangeYuhang Liu, Zhen Zhang, Dong Gong, Mingming Gong et al.ICLR 2024 · 21 citations
