CellPLM: Pre-training of Cell Language Model Beyond Single Cells
Hongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding, Wei Jin, Yuying Xie, Jiliang Tang
Abstract
The current state-of-the-art single-cell pre-trained models are greatly inspired by the success of large language models. They trained transformers by treating genes as tokens and cells as sentences. However, three fundamental differences between single-cell data and natural language data are overlooked: (1) scRNA-seq data are presented as bag-of-genes instead of sequences of RNAs; (2) Cell-cell relations are more intricate and important than inter-sentence relations; and (3) The quantity of single-cell data is considerably inferior to text data, and they are very noisy. In light of these characteristics, we propose a new pre-trained model CellPLM, which takes cells as tokens and tissues as sentences. In addition, we leverage spatially-resolved transcriptomic data in pre-training to facilitate learning cell-cell relationships and introduce a Gaussian mixture prior distribution as an additional inductive bias to overcome data limitation. CellPLM is the first single-cell pre-trained transformer that encodes cell-cell relations and it consistently outperforms existing pre-trained and non-pre-trained models in diverse downstream tasks, with 100 times higher inference speed on generating cell embeddings than previous pre-trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40917df1-8f13-4424-9c49-64cd8ec793d1Cited by top-tier papers15
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang et al.EMNLP 2024 · 28 citations
- HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics DataHiren Madhu, João Felipe Rocha, Tinglin Huang, Siddharth Viswanath et al.ICLR 2026 · 19 citations
- CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell AnalysisYihang Xiao, Jinyi Liu, Yan Zheng, Shaoqing Jiao et al.ICLR 2026 · 11 citations
- A Survey on Foundation Language Models for Single-cell BiologyFan Zhang, Hao Chen, Zhihong Zhu, Ziheng Zhang et al.ACL 2025 · 10 citations
- Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell TranscriptomicsJiayuan Ding, Jianhui Lin, Shiyu Jiang, Yixin Wang et al.NeurIPS 2025 · 4 citations
Builds on4
- Deep Clustering by Gaussian Mixture Variational Autoencoders With Graph EmbeddingLinxiao Yang, Ngai-Man Cheung, Jiaying Li, Jun FangICCV 2019 · 149 citations
- Flowformer: Linearizing Transformers with Conservation FlowsHaixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang et al.ICML 2022 · 130 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq DataJing Gong, Minsheng Hao, Xingyi Cheng, Xin Zeng et al.NeurIPS 2023 · 48 citations
Related papers
- Cell2Sentence: Teaching Large Language Models the Language of BiologyDaniel LeVine, Syed Asad Rizvi, Sacha Lévy, Nazreen Pallikkavaliyaveetil et al.ICML 2024 · 70 citations
- LangCell: Language-Cell Pre-training for Cell Identity UnderstandingSuyuan Zhao, Jiahuan Zhang, Yushuai Wu, Yizhen Luo et al.ICML 2024 · 32 citations
- sciLaMA: A Single-Cell Representation Learning Framework to Leverage Prior Knowledge from Large Language ModelsHongru Hu, Shuwen Zhang, Yongin Choi, Venkat S. Malladi et al.ICML 2025
- SToFM: a Multi-scale Foundation Model for Spatial TranscriptomicsSuyuan Zhao, Yizhen Luo, Ganbo Yang, Yan Zhong et al.ICML 2025
- Generalized Cell Type Annotation and Discovery for Single-Cell RNA-Seq DataYuyao Zhai, Liang Chen, Minghua DengAAAI 2023 · 6 citations
