Lune

NeurIPS2024

A benchmark for prediction of transcriptomic responses to chemical perturbations across cell types

Artur Szalata, Andrew Benz, Robrecht Cannoodt, Mauricio Cortes, Jason Fong, Sunil Kuppasani, Richard Lieberman, Tianyu Liu, Javier Mas-Rosario, Rico Meinl, Jalil Nourisa, Jared Tumiel, Tin M. Tunjic, Mengbo Wang, Noah Weber, Hongyu Zhao, Benedict Anchang, Fabian J. Theis, Malte Luecken, Daniel Burkhardt

2024Year

Abstract

Single-cell transcriptomics has revolutionized our understanding of cellular heterogeneity and drug perturbation effects. However, its high cost and the vast chemical space of potential drugs present barriers to experimentally characterizing the effect of chemical perturbations in all the myriad cell types of the human body. To overcome these limitations, several groups have proposed using machine learning methods to directly predict the effect of chemical perturbations either across cell contexts or chemical space. However, advances in this field have been hindered by a lack of well-designed evaluation datasets and benchmarks. To drive innovation in perturbation modeling, the Open Problems Perturbation Prediction (OP3) benchmark introduces a framework for predicting the effects of small molecule perturbations on cell type-specific gene expression. OP3 leverages the Open Problems in Single-cell Analysis benchmarking infrastructure and is enabled by a new singlecell perturbation dataset, encompassing 146 compounds tested on human blood cells. The benchmark includes diverse data representations, evaluation metrics, and winning methods from our "Single-cell perturbation prediction: generalizing experimental interventions to unseen contexts" competition at NeurIPS 2023. We envision that the OP3 benchmark and competition will drive innovation in single-cell perturbation prediction by improving the accessibility, visibility, and feasibility of this challenge, thereby promoting the impact of machine learning in drug discovery. A living benchmark for perturbation prediction To drive innovation in algorithm development for single-cell perturbation analysis, we set up the OP3 benchmark, including a formalized task definition, a fit-for-purpose benchmarking dataset, and computational infrastructure to support continuously-updated, community-driven benchmarking (Figure 1a ). We outline these features below. Task overview Chemical perturbations induce cell type-specific gene expression changes by interacting with target proteins and altering cellular processes. For example, tamoxifen, a breast cancer drug, binds the estrogen receptor and inhibits cell growth, thereby acting selectively on cells expressing the estrogen receptor [24] . However, the lack of knowledge about mechanisms of action for most compounds hinders predicting their effects on specific cell types. The goal of this task is to leverage data about chemical perturbations in some cell types to infer their impact on gene expression in other cell types. The data is a tensor with three axes: compounds, cell types, and genes. Each value in this tensor is a measurement of the impact on gene expression observed in a specific cell type under a specific chemical perturbation (Section 3.3). Models are provided with the changes in gene expression for all cell types for a subset of compounds. The remaining compounds comprise the test set. These compounds have their differential expression values masked for all genes for a subset of the cell types. The target of this task is to predict these masked differential expression values (Figure 1b ). Generating a single-cell perturbation benchmarking dataset Considerations for data set generation We identified the following properties of an ideal dataset for benchmarking small molecule perturbation prediction: 1. Disease-relevance: To reflect the downstream application to drug discovery, an ideal dataset ought to focus on a disease-relevant biological system. 2. Balanced cellular heterogeneity: Cell types must exhibit distinct perturbation responses but be similar enough that translating compounds' effects is tractable. Diverse perturbations: The compounds should perturb a range of biochemical pathways. 4. Replicates across multiple donors: Capturing perturbation effects across multiple donors enables identifying effects that are preserved across diverse donors. 5. Positive and negative controls: Because of the high degree of technical and biological variability in gene expression measurements, positive and negative controls are essential to accurately estimate the variation attributable to perturbation effects. Open access & informed consent: To ensure open access to benchmarking data collected from human donors, samples must be collected under IRB supervision. This ensures donors give informed consent for public sharing of any derived data. Dataset overview We generated a novel scRNA-seq dataset profiling 146 compounds in PBMCs to provide a high-quality reference benchmark dataset for single-cell perturbation prediction (Figure 1c ). We also included multiome single-nucleus RNA and chromatin accessibility measurements at baseline to facilitate gene regulatory network inference. This effort represents, to date, the largest drug perturbation dataset on primary human tissue with donor replicates [15] , and was specifically designed to satisfy all the criteria above. First, PBMCs comprise an important subset of the human immune