Cross-Modal Translation and Alignment for Survival Analysis
Fengtao Zhou, Hao Chen
Abstract
With the rapid advances in high-throughput sequencing technologies, the focus of survival analysis has shifted from examining clinical indicators to incorporating genomic profiles with pathological images. However, existing methods either directly adopt a straightforward fusion of pathological features and genomic profiles for survival prediction, or take genomic profiles as guidance to integrate the features of pathological images. The former would overlook intrinsic cross-modal correlations. The latter would discard pathological information irrelevant to gene expression. To address these issues, we present a Cross-Modal Translation and Alignment (CMTA) framework to explore the intrinsic cross-modal correlations and transfer potential complementary information. Specifically, we construct two parallel encoder-decoder structures for multi-modal data to integrate intra-modal information and generate cross-modal representation. Taking the generated cross-modal representation to enhance and recalibrate intra-modal representation can significantly improve its discrimination for comprehensive survival analysis. To explore the intrinsic cross-modal correlations, we further design a cross-modal attention module as the information bridge between different modalities to perform cross-modal interactions and transfer complementary information. Our extensive experiments on five public TCGA datasets demonstrate that our proposed framework outperforms the state-of-the-art methods. The source code has been released †.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03e85435-0d15-4562-a494-af3ce4189624Cited by top-tier papers23
- Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyWenhao Tang, Fengtao Zhou, Sheng Huang, Xiang Zhu et al.CVPR 2024 · 70 citations
- Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival PredictionYilan Zhang, Yingxue Xu, Jianqi Chen, Fengying Xie et al.ICLR 2024 · 63 citations
- Graph Domain Adaptation With Dual-Branch Encoder and Two-Level Alignment for Whole Slide Image-Based Survival PredictionYuntao Shou, Xiangyong Cao, Peiqiang Yan, Qiaohui et al.ICCV 2025 · 3 citations
- Act Like a Pathologist: Tissue-Aware Whole Slide Image ReasoningWentao Huang, Weimin Lyu, Peiliang Lou, Qingqiao Hu et al.CVPR 2026 · 3 citations
- PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival PredictionManahil Raza, Ayesha Azam, Talha Qaiser, Nasir M. RajpootICCV 2025 · 2 citations
Builds on3
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen et al.ICCV 2021 · 369 citations
Related papers
- Cancer Survival Prediction by Cyclic Generation and Multi-grained AlignmentYongqi Bu, Qinggang Niu, Zhen Li, Yanyu Xu et al.AAAI 2026
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 80 citations
- MDCS-MoAME: Multi-directional Composite Scanning with Mixture of Attention and Mamba Experts for Cancer Survival PredictionLinjie Qu, Jin Xiao, Xiangrong Liu, Changming Sun et al.CVPR 2026
- PREMISE: Individual Preference-aware Multi-modal Cooperation for Survival PredictionJiaqi Cui, Yilun Li, Xi Wu, Jiliu Zhou et al.ACM MM 2025 · 2 citations
- CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic PredictionYajun An, Jiale Chen, Huan Lin, Zhenbing Liu et al.AAAI 2025 · 1 citation
