Modelling Cellular Perturbations with the Sparse Additive Mechanism Shift Variational Autoencoder
Michael Bereket, Theofanis Karaletsos
Abstract
Generative models of observations under interventions have been a vibrant topic of interest across machine learning and the sciences in recent years. For example, in drug discovery, there is a need to model the effects of diverse interventions on cells in order to characterize unknown biological mechanisms of action. We propose the Sparse Additive Mechanism Shift Variational Autoencoder, SAMS-VAE, to combine compositionality, disentanglement, and interpretability for perturbation models. SAMS-VAE models the latent state of a perturbed sample as the sum of a local latent variable capturing sample-specific variation and sparse global variables of latent intervention effects. Crucially, SAMS-VAE sparsifies these global latent variables for individual perturbations to identify disentangled, perturbation-specific latent subspaces that are flexibly composable. We evaluate SAMS-VAE both quantitatively and qualitatively on a range of tasks using two popular single cell sequencing datasets. In order to measure perturbation-specific model-properties, we also introduce a framework for evaluation of perturbation models based on average treatment effects with links to posterior predictive checks. SAMS-VAE outperforms comparable models in terms of generalization across in-distribution and out-of-distribution tasks, including a combinatorial reasoning task under resource paucity, and yields interpretable latent structures which correlate strongly to known biological mechanisms. Our results suggest SAMS-VAE is an interesting addition to the modeling toolkit for machine learning-driven scientific discovery. * Research supporting this publication conducted while authors were employed at insitro 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Learning Identifiable Factorized Causal Representations of Cellular ResponsesHaiyi Mao, Romain Lopez, Kai Liu, Jan-Christian Huetter et al.NeurIPS 2024 · 10 citations
- Scalable Single-Cell Gene Expression Generation with Latent Diffusion ModelsGiovanni Palla, Sudarshan Babu, Payam Dibaeinia, James Pearce et al.ICML 2026 · 7 citations
- Cradle-VAE: Enhancing Single-Cell Gene Perturbation Modeling with Counterfactual Reasoning-based Artifact DisentanglementSeungheun Baek, Soyon Park, Yan Ting Chok, Junhyun Lee et al.AAAI 2025 · 5 citations
- Doloris: Dual Conditional Diffusion Implicit Bridges with Sparsity Masking Strategy for Unpaired Single-Cell Perturbation EstimationChangxi Chi, Jun Xia, Yufei Huang, Zhuoli Ouyang et al.ICLR 2026 · 4 citations
- Statistical and structural identifiability in representation learningWalter Nelson, Marco Fumero, Theofanis Karaletsos, Francesco LocatelloICLR 2026 · 4 citations
Related papers
- What Makes a Representation Good for Single-Cell Perturbation Prediction?Wenkang Jiang, Yuhang Liu, Yichao Cai, Erdun Gao et al.ICML 2026 · 2 citations
- Interpretable Causal Representation Learning for Biological Data in the Pathway SpaceJesus de la Fuente Cedeño, Robert Lehmann, Carlos Ruiz-Arenas, Jan Voges et al.ICLR 2025
- Identifiability Guarantees for Causal Disentanglement from Soft InterventionsJiaqi Zhang, Kristjan H. Greenewald, Chandler Squires, Akash Srivastava et al.NeurIPS 2023 · 120 citations
- Generative Intervention Models for Causal Perturbation ModelingNora Schneider, Lars Lorch, Niki Kilbertus, Bernhard Schölkopf et al.ICML 2025
- ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse AutoencodersXiangyu Liu, Haodi Lei, Yi Liu, Yang Liu et al.AAAI 2026 · 2 citations
