A Simple and Comprehensive Benchmark for Single-Cell Transcriptomics
Jiaxin Qi, Yan Cui, Kailei Guo, Xiaomin Zhang, Jianqiang Huang, Gaogang Xie
Abstract
Single-cell transcriptomics describes complex molecular features at the individual cell level, serving various roles in biological research, such as enhancing gene expression and predicting drug responses. Due to transcriptomic data structurally resembling sequential data, many researchers have trained numerous transformers on extensive transcriptomic datasets. However, they have consistently neglected to explore the intrinsic properties of the data and the appropriateness of their chosen model architecture. In this paper, we carefully investigate the nature of transcriptomics, identifying three overlooked problems: 1) long-tailed data problem, 2) model selection problem, and 3) evaluation problem. Consequently, by applying the weighted sampling strategy, we address the long-tailed data problem and achieve consistent improvement across all settings. By adapting different model structures to transcriptomic data, we discover that transformers are not the only option. By developing three downstream tasks and fair evaluation metrics, we establish a simple and comprehensive benchmark to validate the effectiveness of models for transcriptomics. Through extensive experiments, we clarify the misunderstandings in the traditional methods and provide competitive baselines, thereby paving the way for future research in this field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- CellPLM: Pre-training of Cell Language Model Beyond Single CellsHongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding et al.ICLR 2024 · 76 citations
- Improving Calibration for Long-Tailed RecognitionZhisheng Zhong, Jiequan Cui, Shu Liu, Jiaya JiaCVPR 2021
Related papers
- Gene Incremental Learning for Single-Cell TranscriptomicsJiaxin Qi, Yan Cui, Jianqiang Huang, Gaogang XieAAAI 2026
- LLM4Cell: Taxonomy and Evaluation of LLM and Agentic Models for Single-Cell BiologySajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo et al.ACL 2026
- A Survey on Foundation Language Models for Single-cell BiologyFan Zhang, Hao Chen, Zhihong Zhu, Ziheng Zhang et al.ACL 2025 · 10 citations
- PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-CancerXiaoshui Huang, Tianlin Zhu, Yifan Zuo, Xue Xia et al.AAAI 2026
- Cell ontology guided transcriptome foundation modelXinyu Yuan, Zhihao Zhan, Zuobai Zhang, Manqi Zhou et al.NeurIPS 2024 · 23 citations
