SToFM: a Multi-scale Foundation Model for Spatial Transcriptomics
Suyuan Zhao, Yizhen Luo, Ganbo Yang, Yan Zhong, Hao Zhou, Zaiqing Nie
Abstract
Spatial Transcriptomics (ST) technologies provide biologists with rich insights into single-cell biology by preserving spatial context of cells. Building foundational models for ST can significantly enhance the analysis of vast and complex data sources, unlocking new perspectives on the intricacies of biological tissues. However, modeling ST data is inherently challenging due to the need to extract multi-scale information from tissue slices containing vast numbers of cells. This process requires integrating macro-scale tissue morphology, micro-scale cellular microenvironment, and gene-scale gene expression profile. To address this challenge, we propose SToFM, a multi-scale Spatial Transcriptomics Foundation Model. SToFM first performs multi-scale information extraction on each ST slice, to construct a set of ST sub-slices that aggregate macro-, micro-and gene-scale information. Then an SE(2) Transformer is used to obtain high-quality cell representations from the sub-slices. Additionally, we construct SToCorpus-88M, the largest high-resolution spatial transcriptomics corpus for pretraining. SToFM achieves outstanding performance on a variety of downstream tasks, such as tissue region semantic segmentation and cell type annotation, demonstrating its comprehensive understanding of ST data through capturing and integrating multi-scale information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abb8455b-dd5b-4f4a-a246-ce9ea86da801Cited by top-tier papers2
- Adapting a Pre-trained Single-Cell Foundation Model to Spatial Gene Expression Generation from Histology ImagesDonghai Fang, Yongheng Li, Zhen WANG, Yuansong Zeng et al.CVPR 2026 · 2 citations
- SpaEF: Spatially Resolved Transcriptomics Data Element-Wise Denoising Framework Powered by Large ModelsZekuan Shang, Xiaosong Han, Liupu Wang, Wei Du et al.ICML 2026
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- CellPLM: Pre-training of Cell Language Model Beyond Single CellsHongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding et al.ICLR 2024 · 76 citations
Related papers
- SPATIA: Multimodal Generation and Prediction of Spatial Cell PhenotypesZhenglun Kong, Mufan Qiu, John Boesen, xiang lin et al.ICML 2026 · 1 citation
- HiST: A Hierarchical Sparse Transformer for Cross-Modal Spatial Transcriptomics ModelingWeiyi Wu, Xinwen Xu, Xingjian Diao, Siting Li et al.ICML 2026
- HiFusion: Hierarchical Intra-Spot Alignment and Regional Context Fusion for Spatial Gene Expression Prediction from HistopathologyZiqiao Weng, Yaoyu Fang, Jiahe Qian, Xinkun Wang et al.AAAI 2026
- HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics DataHiren Madhu, João Felipe Rocha, Tinglin Huang, Siddharth Viswanath et al.ICLR 2026 · 19 citations
- Predicting Spatial Transcriptomics from Histology Images via High-Order Multi-Cell Interaction ModelingYouhan Sun, Jiahua Rao, Kangrui Du, Jiancong Xie et al.CVPR 2026
