A Survey on Foundation Language Models for Single-cell Biology
Fan Zhang, Hao Chen, Zhihong Zhu, Ziheng Zhang, Zhenxi Lin, Ziyue Qiao, Yefeng Zheng, Xian Wu
摘要
The recent advancements in language models have significantly catalyzed progress in computational biology. A growing body of research strives to construct unified foundation models for single-cell biology, with language models serving as the cornerstone. In this paper, we systematically review the developments in foundation language models designed specifically for single-cell biology. Our survey offers a thorough analysis of various incarnations of single-cell foundation language models, viewed through the lens of both pre-trained language models (PLMs) and large language models (LLMs). This includes an exploration of data tokenization strategies, pre-training/tuning paradigms, and downstream single-cell data analysis tasks. Additionally, we discuss the current challenges faced by these pioneering works and speculate on future research directions. Overall, this survey provides a comprehensive overview of the existing single-cell foundation language models, paving the way for future research endeavors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information DecodingZhihong Zhu, Yunyan Zhang, Fan Zhang, Bowen Xing 等AAAI 2026 · 被引用 1 次
- scCluBench: Comprehensive Benchmarking of Clustering Algorithms for Single-Cell RNA SequencingPing Xu, Zaitian Wang, Zhirui Wang, Pengjiang Li 等AAAI 2026
- (Be Cautious!) Bio-Foundation Models Are Not Yet Robust to Biologically Plausible Perturbations and ML TransformationsJinhao Duan, Ruichen Zhang, Gengwei Zhang, Huaizhi Qu 等ICML 2026
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi 等SIGIR 2024 · 被引用 182 次
相关 Paper
- LLM4Cell: Taxonomy and Evaluation of LLM and Agentic Models for Single-Cell BiologySajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo 等ACL 2026
- Cell2Sentence: Teaching Large Language Models the Language of BiologyDaniel LeVine, Syed Asad Rizvi, Sacha Lévy, Nazreen Pallikkavaliyaveetil 等ICML 2024 · 被引用 70 次
- LangCell: Language-Cell Pre-training for Cell Identity UnderstandingSuyuan Zhao, Jiahuan Zhang, Yushuai Wu, Yizhen Luo 等ICML 2024 · 被引用 32 次
- CellPLM: Pre-training of Cell Language Model Beyond Single CellsHongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding 等ICLR 2024 · 被引用 76 次
- SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented EvaluationJiahao Zhao, Feng Jiang, Shaowei Qin, Zhonghui Zhang 等ICLR 2026 · 被引用 4 次
