What Makes a Representation Good for Single-Cell Perturbation Prediction?
Wenkang Jiang, Yuhang Liu, Yichao Cai, Erdun Gao, Jiayi Dong, Ehsan Abbasnejad, Lina Yao, Javen Qinfeng Shi
Abstract
Single-cell perturbation modeling is fundamental for understanding and predicting cellular responses to genetic perturbations. However, existing approaches, from causal representation learning to foundation models, often struggle with an overlooked challenge: gene expression is dominated by perturbation-invariant information, while perturbation-specific signals are intrinsically sparse. As a result, learned representations either entangle invariant and perturbation-specific information, leading to spurious and non-generalizable predictors, or suppress perturbation-specific signals altogether, rendering them ineffective for prediction. To address this, we propose PerturbedVAE, a general framework designed to resolve this signal imbalance. The framework explicitly separates perturbation-specific information from dominant invariant structure and recovers causal representations to effectively utilize such information for prediction. We further provide an identifiability analysis that characterizes the conditions under which sparse perturbation effects can be reliably recovered, thereby clarifying how the framework can be concretely specified under such conditions. Empirically, PerturbedVAE achieves state-of-the-art performance on a widely used benchmark across multiple evaluation settings, yielding significant gains on out-of-distribution combinatorial predictions and uncovering interpretable perturbation-response programs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- Interventional Causal Representation LearningKartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua BengioICML 2023 · 143 citations
- CITRIS: Causal Identifiability from Temporal Intervened SequencesPhillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano et al.ICML 2022 · 136 citations
Related papers
- Cradle-VAE: Enhancing Single-Cell Gene Perturbation Modeling with Counterfactual Reasoning-based Artifact DisentanglementSeungheun Baek, Soyon Park, Yan Ting Chok, Junhyun Lee et al.AAAI 2025 · 5 citations
- Beyond Independent Genes: Learning Module-Inductive Representations for Single-Cell Gene Perturbation PredictionJiafa Ruan, Ruijie Quan, Liyang Xu, Zongxin Yang et al.ICML 2026 · 3 citations
- Modelling Cellular Perturbations with the Sparse Additive Mechanism Shift Variational AutoencoderMichael Bereket, Theofanis KaraletsosNeurIPS 2023 · 59 citations
- scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation PredictionChenglei Yu, Chuanrui Wang, Bangyan Liao, Tailin WuICLR 2026 · 15 citations
- Learning Cross-Domain Representations for Transferable Drug Perturbations on Single-Cell Transcriptional ResponsesHui Liu, Shikai JinAAAI 2025 · 1 citation
