CellPLM: Pre-training of Cell Language Model Beyond Single Cells
Hongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding, Wei Jin, Yuying Xie, Jiliang Tang
摘要
The current state-of-the-art single-cell pre-trained models are greatly inspired by the success of large language models. They trained transformers by treating genes as tokens and cells as sentences. However, three fundamental differences between single-cell data and natural language data are overlooked: (1) scRNA-seq data are presented as bag-of-genes instead of sequences of RNAs; (2) Cell-cell relations are more intricate and important than inter-sentence relations; and (3) The quantity of single-cell data is considerably inferior to text data, and they are very noisy. In light of these characteristics, we propose a new pre-trained model CellPLM, which takes cells as tokens and tissues as sentences. In addition, we leverage spatially-resolved transcriptomic data in pre-training to facilitate learning cell-cell relationships and introduce a Gaussian mixture prior distribution as an additional inductive bias to overcome data limitation. CellPLM is the first single-cell pre-trained transformer that encodes cell-cell relations and it consistently outperforms existing pre-trained and non-pre-trained models in diverse downstream tasks, with 100 times higher inference speed on generating cell embeddings than previous pre-trained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics DataHiren Madhu, João Felipe Rocha, Tinglin Huang, Siddharth Viswanath 等ICLR 2026 · 被引用 19 次
- CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell AnalysisYihang Xiao, Jinyi Liu, Yan Zheng, Shaoqing Jiao 等ICLR 2026 · 被引用 11 次
- A Survey on Foundation Language Models for Single-cell BiologyFan Zhang, Hao Chen, Zhihong Zhu, Ziheng Zhang 等ACL 2025 · 被引用 10 次
- Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell TranscriptomicsJiayuan Ding, Jianhui Lin, Shiyu Jiang, Yixin Wang 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper4
- Deep Clustering by Gaussian Mixture Variational Autoencoders With Graph EmbeddingLinxiao Yang, Ngai-Man Cheung, Jiaying Li, Jun FangICCV 2019 · 被引用 149 次
- Flowformer: Linearizing Transformers with Conservation FlowsHaixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang 等ICML 2022 · 被引用 130 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq DataJing Gong, Minsheng Hao, Xingyi Cheng, Xin Zeng 等NeurIPS 2023 · 被引用 48 次
相关 Paper
- Cell2Sentence: Teaching Large Language Models the Language of BiologyDaniel LeVine, Syed Asad Rizvi, Sacha Lévy, Nazreen Pallikkavaliyaveetil 等ICML 2024 · 被引用 70 次
- LangCell: Language-Cell Pre-training for Cell Identity UnderstandingSuyuan Zhao, Jiahuan Zhang, Yushuai Wu, Yizhen Luo 等ICML 2024 · 被引用 32 次
- sciLaMA: A Single-Cell Representation Learning Framework to Leverage Prior Knowledge from Large Language ModelsHongru Hu, Shuwen Zhang, Yongin Choi, Venkat S. Malladi 等ICML 2025
- SToFM: a Multi-scale Foundation Model for Spatial TranscriptomicsSuyuan Zhao, Yizhen Luo, Ganbo Yang, Yan Zhong 等ICML 2025
- Generalized Cell Type Annotation and Discovery for Single-Cell RNA-Seq DataYuyao Zhai, Liang Chen, Minghua DengAAAI 2023 · 被引用 6 次
