Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell Transcriptomics
Jiayuan Ding, Jianhui Lin, Shiyu Jiang, Yixin Wang, Ziyang Miao, Zhaoyu Fang, Jiliang Tang, Min Li, Xiaojie Qiu
Abstract
Foundation models (FMs) have shown great promise in single-cell genomics, yet current approaches, such as scGPT, Geneformer, and scFoundation, rely on centralized training and language modeling objectives that overlook the tabular nature of single-cell data and raise significant privacy concerns. We present T ABULA , a foundation model designed for single-cell transcriptomics, which integrates a novel tabular modeling objective and federated learning framework to enable privacy-preserving pretraining across decentralized datasets. T ABULA directly models the cell-by-gene expression matrix through column-wise gene reconstruction and row-wise cell contrastive learning, capturing both gene-level relationships and cell-level heterogeneity without imposing artificial gene sequence order. Extensive experiments demonstrate the effectiveness of T ABULA : despite using only half the pretraining data, T ABULA achieves state-of-the-art performance across key tasks, including gene imputation, perturbation prediction, cell type annotation, and multi-omics integration. It is important to note that as public single-cell datasets continue to grow, T ABULA provides a scalable and privacy-aware foundation that not only validates the feasibility of federated tabular modeling, but also establishes a generalizable framework for training future models under similar privacy-preserving settings. All resources are openly available at https://github.com/aristoteleo/tabula to support broad community adoption and future methodological advances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86f29d41-b3fb-4058-80cf-c42f92196f0aBuilds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li et al.ICML 2023 · 111 citations
- CellPLM: Pre-training of Cell Language Model Beyond Single CellsHongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding et al.ICLR 2024 · 76 citations
Related papers
- Cell ontology guided transcriptome foundation modelXinyu Yuan, Zhihao Zhan, Zuobai Zhang, Manqi Zhou et al.NeurIPS 2024 · 23 citations
- Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation ModelsJiaxin Qi, Hang Li, Yan Cui, Yuhua Zheng et al.ICML 2026
- scDEBART: Predicting in silico Single-Cell Perturbation Responses via Large-Scale Differential Expression LearningJieun Sung, Wankyu KimICML 2026
- PertEval-scFM: Benchmarking Single-Cell Foundation Models for Perturbation Effect PredictionAaron Wenteler, Martina Occhetta, Nikhil Branson, Victor Curean et al.ICML 2025
- ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and ExpressionMingxuan Wang, Gaoyang Jiang, ZiJia Ren, Cheng Chen et al.ICML 2026
