Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
Minghao Han, Dingkang Yang, Linhao Qu, Zizhi Chen, Gang Li, Han Wang, Jiacong Wang, Lihua Zhang
摘要
Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottlenecks. In this paper, we propose STAMP, a Spatial Transcriptomics-Augmented Multimodal Pathology representation learning framework that integrates spatially-resolved gene expression profiles to enable molecule-guided joint embedding of pathology images and transcriptomic data. Our study shows that self-supervised, gene-guided training provides a robust and task-agnostic signal for learning pathology image representations. Incorporating spatial context and multi-scale information further enhances model performance and generalizability. To support this, we constructed SpaVis-6M, the largest Visium-based spatial transcriptomics dataset to date, and trained a spatially-aware gene encoder on this resource. Leveraging hierarchical multi-scale contrastive alignment and cross-scale patch localization mechanisms, STAMP effectively aligns spatial transcriptomics with pathology images, capturing spatial structure and molecular variation. We validate STAMP across six datasets and four downstream tasks, where it consistently achieves strong performance. These results highlight the value and necessity of integrating spatially resolved molecular supervision for advancing multimodal learning in computational pathology. The code is included in the supplementary materials. The pretrained weights and SpaVis-6M are available at: https://github.com/Hanminghao/STAMP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype ControlMinghao Han, Yichen Liu, Yizhou Liu, Zizhi Chen 等CVPR 2026 · 被引用 5 次
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation ModelsZizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Spatially Resolved Gene Expression Prediction from Histology Images via Bi-modal Contrastive LearningRonald Xie, Kuan Pang, Sai Chung, Catia Perciani 等NeurIPS 2023 · 被引用 125 次
相关 Paper
- HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics PredictionChen Zhang, Yilu An, Ying Chen, Hao Li 等CVPR 2026
- ST-LLM: Spatial Transcriptomics Embedding with Large Language ModelsZhetao Xu, Xiaohua Wan, Le Li, Shuang Feng 等AAAI 2026
- Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression InferenceZhiceng Shi, Changmiao Wang, Jun Wan, Wenwen MinCVPR 2026 · 被引用 1 次
- Accurate Spatial Gene Expression Prediction by Integrating Multi-Resolution FeaturesYoungmin Chung, Ji Hun Ha, Kyeong Chan Im, Joo Sang LeeCVPR 2024
- When Genes Speak: A Semantic-Guided Framework for Spatially Resolved Transcriptomics Data ClusteringJiangkai Long, Yanran Zhu, Chang Tang, Kun Sun 等AAAI 2026
