Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell Transcriptomics
Jiayuan Ding, Jianhui Lin, Shiyu Jiang, Yixin Wang, Ziyang Miao, Zhaoyu Fang, Jiliang Tang, Min Li, Xiaojie Qiu
摘要
Foundation models (FMs) have shown great promise in single-cell genomics, yet current approaches, such as scGPT, Geneformer, and scFoundation, rely on centralized training and language modeling objectives that overlook the tabular nature of single-cell data and raise significant privacy concerns. We present T ABULA , a foundation model designed for single-cell transcriptomics, which integrates a novel tabular modeling objective and federated learning framework to enable privacy-preserving pretraining across decentralized datasets. T ABULA directly models the cell-by-gene expression matrix through column-wise gene reconstruction and row-wise cell contrastive learning, capturing both gene-level relationships and cell-level heterogeneity without imposing artificial gene sequence order. Extensive experiments demonstrate the effectiveness of T ABULA : despite using only half the pretraining data, T ABULA achieves state-of-the-art performance across key tasks, including gene imputation, perturbation prediction, cell type annotation, and multi-omics integration. It is important to note that as public single-cell datasets continue to grow, T ABULA provides a scalable and privacy-aware foundation that not only validates the feasibility of federated tabular modeling, but also establishes a generalizable framework for training future models under similar privacy-preserving settings. All resources are openly available at https://github.com/aristoteleo/tabula to support broad community adoption and future methodological advances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li 等ICML 2023 · 被引用 111 次
- CellPLM: Pre-training of Cell Language Model Beyond Single CellsHongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding 等ICLR 2024 · 被引用 76 次
相关 Paper
- Cell ontology guided transcriptome foundation modelXinyu Yuan, Zhihao Zhan, Zuobai Zhang, Manqi Zhou 等NeurIPS 2024 · 被引用 23 次
- Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation ModelsJiaxin Qi, Hang Li, Yan Cui, Yuhua Zheng 等ICML 2026
- scDEBART: Predicting in silico Single-Cell Perturbation Responses via Large-Scale Differential Expression LearningJieun Sung, Wankyu KimICML 2026
- PertEval-scFM: Benchmarking Single-Cell Foundation Models for Perturbation Effect PredictionAaron Wenteler, Martina Occhetta, Nikhil Branson, Victor Curean 等ICML 2025
- ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and ExpressionMingxuan Wang, Gaoyang Jiang, ZiJia Ren, Cheng Chen 等ICML 2026
