Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images
Richard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen, Drew F. K. Williamson, Trevor Manz, Maha Shady, Faisal Mahmood
Abstract
Survival outcome prediction is a challenging weakly-supervised and ordinal regression task in computational pathology that involves modeling complex interactions within the tumor microenvironment in gigapixel whole slide images (WSIs). Despite recent progress in formulating WSIs as bags for multiple instance learning (MIL), representation learning of entire WSIs remains an open and challenging problem, especially in overcoming: 1) the computational complexity of feature aggregation in large bags, and 2) the data heterogeneity gap in incorporating biological priors such as genomic measurements. In this work, we present a Multimodal Co-Attention Transformer (MCAT) framework that learns an interpretable, dense co-attention mapping between WSIs and genomic features formulated in an embedding space. Inspired by approaches in Visual Question Answering (VQA) that can attribute how word embed-dings attend to salient objects in an image when answering a question, MCAT learns how histology patches attend to genes when predicting patient survival. In addition to visualizing multimodal interactions, our co-attention trans-formation also reduces the space complexity of WSI bags, which enables the adaptation of Transformer layers as a general encoder backbone in MIL. We apply our proposed method on five different cancer datasets (4,730 WSIs, 67 million patches). Our experimental results demonstrate that the proposed method consistently achieves superior performance compared to the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 817a23aa-0ef4-4035-a9f7-b6d97ba65e62Cited by top-tier papers45
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionYingxue Xu, Hao ChenICCV 2023 · 132 citations
- Cross-Modal Translation and Alignment for Survival AnalysisFengtao Zhou, Hao ChenICCV 2023 · 123 citations
- Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkPaul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling et al.NeurIPS 2023 · 120 citations
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 80 citations
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- CAMEL: A Weakly Supervised Learning Framework for Histopathology Image SegmentationGang Xu, Zhigang Song, Zhuo Sun, Calvin Ku et al.ICCV 2019 · 187 citations
- HistoSegNet: Semantic Segmentation of Histological Tissue Type in Whole Slide ImagesLyndon Chan, Mahdi S. Hosseini, Corwyn Rowsell, Konstantinos N. Plataniotis et al.ICCV 2019 · 131 citations
- Multiple Instance Captioning: Learning Representations From Histopathology Textbooks and ArticlesJevgenij Gamper, Nasir M. RajpootCVPR 2021
Related papers
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang et al.AAAI 2023 · 44 citations
- Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational PathologyPei Liu, Luping Ji, Jiaxiang Gou, Bo Fu et al.ICLR 2025
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionGuillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson et al.CVPR 2024
- CO-PILOT: Dynamic Top-Down Point Cloud with Conditional Neighborhood Aggregation for Multi-Gigapixel Histopathology Image RepresentationRamin Nakhli, Allen W. Zhang, Ali Khajegili Mirabadi, Katherine Rich et al.ICCV 2023 · 8 citations
