xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data
Jing Gong, Minsheng Hao, Xingyi Cheng, Xin Zeng, Chiming Liu, Jianzhu Ma, Xuegong Zhang, Taifeng Wang, Le Song
摘要
Advances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M records for humans with each record measuring 20,000 genes. This highlights the need for unsupervised representation learning to fully ingest these data, yet classical transformer architectures are prohibitive to train on such data in terms of both computation and memory. To address this challenge, we propose a novel asymmetric encoder-decoder transformer for scRNA-seq data, called xTrimoGene α (or xTrimoGene for short) 4 , which leverages the sparse characteristic of the data to scale up the pre-training. This scalable design of xTrimoGene reduces FLOPs by one to two orders of magnitude compared to classical transformers while maintaining high accuracy, enabling us to train the largest transformer models over the largest scRNA-seq dataset today. Our experiments also show that the performance of xTrimoGene improves as we scale up the model sizes, and it also leads to SOTA performance over various downstream tasks, such as cell type annotation, perturb-seq effect prediction, and drug combination prediction. xTrimoGene model is now available for use as a service via the following link: https://api.biomap.com/xTrimoGene/apply * Equal contribution. Mingsheng Hao conducted this work during his internship at BioMap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- CellPLM: Pre-training of Cell Language Model Beyond Single CellsHongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding 等ICLR 2024 · 被引用 76 次
- scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and DiscoveryYiming Gao, Zhen Wang, Jefferson Chen, Mark Antkowiak 等NeurIPS 2025 · 被引用 8 次
- ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibiltiy DataYifeng Jiao, Yuchen Liu, Yu Zhang, Xin Guo 等NeurIPS 2025 · 被引用 1 次
- CountsDiff: A diffusion model on the natural numbers for generation and imputation of count-based dataRenzo Soatto, Anders Hoel, Greycen Ren, Shorna Alam 等ICML 2026
- Skipping the Zeros in Diffusion Models for Sparse Data GenerationPhil Sidney Ostheimer, Mayank Kumar Nagda, Andriy Balinskyy, Gabriel Rodrigues 等ICML 2026
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 被引用 338 次
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
相关 Paper
- sciLaMA: A Single-Cell Representation Learning Framework to Leverage Prior Knowledge from Large Language ModelsHongru Hu, Shuwen Zhang, Yongin Choi, Venkat S. Malladi 等ICML 2025
- Bipartite Graph Attention-based Clustering for Large-scale scRNA-seq DataZhuomin Liang, Liang Bai, Xian YangICML 2026
- Generalized Cell Type Annotation and Discovery for Single-Cell RNA-Seq DataYuyao Zhai, Liang Chen, Minghua DengAAAI 2023 · 被引用 6 次
- PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-CancerXiaoshui Huang, Tianlin Zhu, Yifan Zuo, Xue Xia 等AAAI 2026
- Gene Regulatory Network Inference using 3D Convolutional Neural NetworkYue Fan, Xiuli MaAAAI 2021 · 被引用 23 次
