Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction
Yingxue Xu, Hao Chen
Abstract
Survival prediction is a complicated ordinal regression task that aims to predict the ranking risk of death, which generally benefits from the integration of histology and genomic data. Despite the progress in joint learning from pathology and genomics, existing methods still suffer from challenging issues: 1) Due to the large size of pathological images, it is difficult to effectively represent the gigapixel whole slide images (WSIs). 2) Interactions within tumor microenvironment (TME) in histology are essential for survival analysis. Although current approaches attempt to model these interactions via co-attention between histology and genomic data, they focus on only dense local similarity across modalities, which fails to capture global consistency between potential structures, i.e. TME-related interactions of histology and co-expression of genomic data. To address these challenges, we propose a Multimodal Optimal Transport-based Co-Attention Transformer framework with global structure consistency, in which optimal transport (OT) is applied to match patches of a WSI and genes embeddings for selecting informative patches to represent the gigapixel WSI. More importantly, OT-based co-attention provides a global awareness to effectively capture structural interactions within TME for survival prediction. To overcome high computational complexity of OT, we propose a robust and efficient implementation over micro-batch of WSI patches by approximating the original OT with unbalanced mini-batch OT. Extensive experiments show the superiority of our method on five benchmark datasets compared to the state-of-the-art methods. The code is released 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 493ec174-21db-4ec4-9e4b-b6f80c82c9c9Cited by top-tier papers35
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 80 citations
- Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival PredictionYilan Zhang, Yingxue Xu, Jianqi Chen, Fengying Xie et al.ICLR 2024 · 63 citations
- Multimodal Prototyping for cancer survival predictionAndrew H. Song, Richard J. Chen, Guillaume Jaume, Anurag J. Vaidya et al.ICML 2024 · 53 citations
- Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyAndrew H. Song, Richard J. Chen, Tong Ding, Drew F. K. Williamson et al.CVPR 2024 · 51 citations
- DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival PredictionYucheng Xing, ling huang, Jingying Ma, Ruping Hong et al.ICML 2026 · 8 citations
Builds on9
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen et al.ICCV 2021 · 369 citations
- CAMEL: A Weakly Supervised Learning Framework for Histopathology Image SegmentationGang Xu, Zhigang Song, Zhuo Sun, Calvin Ku et al.ICCV 2019 · 187 citations
- Unbalanced minibatch Optimal Transport; applications to Domain AdaptationKilian Fatras, Thibault Séjourné, Rémi Flamary, Nicolas CourtyICML 2021 · 183 citations
- OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal TransportZongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He et al.NeurIPS 2022 · 117 citations
Related papers
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionGuillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson et al.CVPR 2024
- Tumor Micro-Environment Interactions Guided Graph Learning for Survival Analysis of Human Cancers from Whole-Slide Pathological ImagesWei Shao, Yangyang Shi, Daoqiang Zhang, Junjie Zhou et al.CVPR 2024
- H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival PredictionJiaqi Yang, Wenting Chen, Xiangjian He, Yuanbai Li et al.CVPR 2026
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang et al.AAAI 2023 · 44 citations
- Sparse Multi-Modal Graph Transformer with Shared-Context Processing for Representation Learning of Giga-pixel ImagesRamin Nakhli, Puria Azadi Moghadam, Haoyang Mi, Hossein Farahani et al.CVPR 2023
