SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region Detection
Shuailin Xue, Jun Wan, Lihua Zhang, Wenwen Min
Abstract
Accurate detection of cancer tissue regions (CTR) enables deeper analysis of the tumor microenvironment and offers crucial insights into treatment response. Traditional CTR detection methods, which typically rely on the rich cellular morphology in histology images, are susceptible to a high rate of false positives due to morphological similarities across different tissue regions. The groundbreaking advances in spatial transcriptomics (ST) provide detailed cellular phenotypes and spatial localization information, offering new opportunities for more accurate cancer region detection. However, current methods are unable to effectively integrate histology images with ST data, especially in the context of cross-sample and cross-platform/batch settings for accomplishing the CTR detection. To address this challenge, we propose SpaCRD, a transfer learning-based method that deeply integrates histology images and ST data to enable reliable CTR detection across diverse samples, platforms, and batches. Once trained on source data, SpaCRD can be readily generalized to accurately detect cancerous regions across samples from different platforms and batches. The core of SpaCRD is a category-regularized variational reconstruction-guided bidirectional cross-attention fusion network, which enables the model to adaptively capture latent co-expression patterns between histological features and gene expression from multiple perspectives. Extensive benchmark analysis on 23 matched histology-ST datasets spanning various disease types, platforms, and batches demonstrates that SpaCRD consistently outperforms existing eight state-of-the-art methods in CTR detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- Learning and Evaluating Representations for Deep One-Class ClassificationKihyuk Sohn, Chun-Liang Li, Jinsung Yoon, Minho Jin et al.ICLR 2021 · 243 citations
- MEATRD: Multimodal Anomalous Tissue Region Detection Enhanced with Spatial TranscriptomicsKaichen Xu, Qilong Wu, Yan Lu, Yinan Zheng et al.AAAI 2025 · 5 citations
Related papers
- Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression InferenceZhiceng Shi, Changmiao Wang, Jun Wan, Wenwen MinCVPR 2026 · 1 citation
- HiFusion: Hierarchical Intra-Spot Alignment and Regional Context Fusion for Spatial Gene Expression Prediction from HistopathologyZiqiao Weng, Yaoyu Fang, Jiahe Qian, Xinkun Wang et al.AAAI 2026
- Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial TranscriptomicsHang Shi, Ruocheng Yang, Wenjie You, Zhilin Huang et al.CVPR 2026
- SPATIA: Multimodal Generation and Prediction of Spatial Cell PhenotypesZhenglun Kong, Mufan Qiu, John Boesen, xiang lin et al.ICML 2026 · 1 citation
- HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics PredictionChen Zhang, Yilu An, Ying Chen, Hao Li et al.CVPR 2026
