Multi-Modal Representation for Spatially Resolved Transcriptomics Based on Global Correlation and Dynamic Cluster Discovery
Chuanxiu Li, Shengwu Xiong, Zhenyu Xiong, Mingxi Sun, Qixiang Zou, Yi Rong
Abstract
Constructing effective representations of spatial resolved transcriptomics (SRT) data, by appropriately characterizing the coherence in gene expression and histology with the spatial information of each sequencing spot, plays an important role in understanding the organization and function of complex tissues. Although much progress has been made, existing SRT representation methods typically establish local associative relationships for each spot only with those located in its surrounding spatial areas, thus failing to capture long-range correlations between distant regions. In addition, the absence of supervision signals on which cluster (with similar biological functions, pathological states or cell types) each spot should belong to also poses a great challenge in deriving an effective representation of SRT data. To this end, we propose a novel Multi-Modal SRT Representation Learning (MMSRL) method based on global spot correlation and dynamic cluster discovery. Specifically, given the gene expression and histological image, MMSRL first builds individual graph convolutional networks (GCNs) for these two modalities and bridges them through the adjacency matrix generated from the spatial locations of different spots. The extracted GCN features are then processed by a correlative self-attention operation to enhance their long-range correlations within each modality. Meanwhile, we also design a multi-modal interaction module (MMIM) to make these two-modal features interact with each other, and align their global correlation information across modalities. After that, we develop an attention-weighted fusion module (AWFM) to adaptively fuse the enhanced intra- and inter-modal features obtained above, so as to effectively integrate the multi-modal information. Finally, a dynamic cluster discovery process, which unsupervisedly assigns cluster labels to each spot, is incorporated to further refine the fused multi-modal representations in a contrastive learning manner.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial TranscriptomicsHang Shi, Ruocheng Yang, Wenjie You, Zhilin Huang et al.CVPR 2026
- Multi-modal Topology-embedded Graph Learning for Spatially Resolved Genes Prediction from Pathology Images with Prior Gene Similarity InformationHang Shi, Changxi Chi, Peng Wan, Daoqiang Zhang et al.CVPR 2025
- MEATRD: Multimodal Anomalous Tissue Region Detection Enhanced with Spatial TranscriptomicsKaichen Xu, Qilong Wu, Yan Lu, Yinan Zheng et al.AAAI 2025 · 5 citations
- Heterogeneous Graph Guided Contrastive Learning for Spatially Resolved Transcriptomics DataXiao He, Chang Tang, Xinwang Liu, Chuankun Li et al.ACM MM 2024 · 9 citations
- GROVER: Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics FusionYongjun Xiao, Dian Meng, Xinlei Huang, Yanran Liu et al.AAAI 2026
