Causal Representation Learning from Multimodal Biomedical Observations
Yuewen Sun, Lingjing Kong, Guangyi Chen, Loka Li, Gongxu Luo, Zijian Li, Yixuan Zhang, Yujia Zheng, Mengyue Yang, Petar Stojanov, Eran Segal, Eric P. Xing, Kun Zhang
Abstract
Prevalent in biomedical applications (e.g., human phenotype research), multimodal datasets can provide valuable insights into the underlying physiological mechanisms. However, current machine learning (ML) models designed to analyze these datasets often lack interpretability and identifiability guarantees, which are essential for biomedical research. Recent advances in causal representation learning have shown promise in identifying interpretable latent causal variables with formal theoretical guarantees. Unfortunately, most current work on multimodal distributions either relies on restrictive parametric assumptions or yields only coarse identification results, limiting their applicability to biomedical research that favors a detailed understanding of the mechanisms. In this work, we aim to develop flexible identification conditions for multimodal data and principled methods to facilitate the understanding of biomedical datasets. Theoretically, we consider a nonparametric latent distribution (c.f., parametric assumptions in previous work) that allows for causal relationships across potentially different modalities. We establish identifiability guarantees for each latent component, extending the subspace identification results from previous work. Our key theoretical contribution is the structural sparsity of causal connections between modalities, which, as we will discuss, is natural for a large collection of biomedical systems. Empirically, we present a practical framework to instantiate our theoretical insights. We demonstrate the effectiveness of our approach through extensive experiments on both numerical and synthetic datasets. Results on a real-world human phenotype dataset are consistent with established biomedical research, validating our theoretical and methodological framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e5d31de-02f9-4840-a1e2-35132d4b0a4dCited by top-tier papers9
- Towards Identifiability of Hierarchical Temporal Causal Representation LearningZijian Li, Minghao Fu, Junxian Huang, Yifan Shen et al.NeurIPS 2025 · 10 citations
- The third pillar of causal analysis? A measurement perspective on causal representationsDingling Yao, Shimeng Huang, Riccardo Cadei, Kun Zhang et al.NeurIPS 2025 · 5 citations
- TRACE: Trajectory Recovery for Continuous Mechanism Evolution in Causal Representation LearningShicheng Fan, Kun Zhang, Lu ChengICML 2026 · 3 citations
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior TraitsLoka Li, Wong Yu Kang, Minghao Fu, Guangyi Chen et al.ICLR 2026 · 2 citations
- ProM3E: Probabilistic Masked MultiModal Embedding Model for EcologySrikumar Sastry, Subash Khanal, Aayush Dhakal, Jiayu Lin et al.CVPR 2026 · 1 citation
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- Contrastive Learning Inverts the Data Generating ProcessRoland S. Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge et al.ICML 2021 · 264 citations
- Interventional Causal Representation LearningKartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua BengioICML 2023 · 143 citations
Related papers
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele et al.NeurIPS 2023 · 127 citations
- Sample-efficient Learning of Concepts with Theoretical Guarantees: from Data to Concepts without InterventionsHidde Fokkema, Tim van Erven, Sara MagliacaneNeurIPS 2025 · 7 citations
- Causal Representation Learning from Multiple Distributions: A General SettingKun Zhang, Shaoan Xie, Ignavier Ng, Yujia ZhengICML 2024 · 61 citations
- CARL: Preserving Causal Structure in Representation LearningYulong Li, Xiwei Liu, Feilong Tang, Zhixiang Lu et al.ICLR 2026
- From Causal to Concept-Based Representation LearningGoutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf et al.NeurIPS 2024 · 37 citations
