Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images
Richard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen, Drew F. K. Williamson, Trevor Manz, Maha Shady, Faisal Mahmood
摘要
Survival outcome prediction is a challenging weakly-supervised and ordinal regression task in computational pathology that involves modeling complex interactions within the tumor microenvironment in gigapixel whole slide images (WSIs). Despite recent progress in formulating WSIs as bags for multiple instance learning (MIL), representation learning of entire WSIs remains an open and challenging problem, especially in overcoming: 1) the computational complexity of feature aggregation in large bags, and 2) the data heterogeneity gap in incorporating biological priors such as genomic measurements. In this work, we present a Multimodal Co-Attention Transformer (MCAT) framework that learns an interpretable, dense co-attention mapping between WSIs and genomic features formulated in an embedding space. Inspired by approaches in Visual Question Answering (VQA) that can attribute how word embed-dings attend to salient objects in an image when answering a question, MCAT learns how histology patches attend to genes when predicting patient survival. In addition to visualizing multimodal interactions, our co-attention trans-formation also reduces the space complexity of WSI bags, which enables the adaptation of Transformer layers as a general encoder backbone in MIL. We apply our proposed method on five different cancer datasets (4,730 WSIs, 67 million patches). Our experimental results demonstrate that the proposed method consistently achieves superior performance compared to the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen 等CVPR 2022 · 被引用 490 次
- Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionYingxue Xu, Hao ChenICCV 2023 · 被引用 132 次
- Cross-Modal Translation and Alignment for Survival AnalysisFengtao Zhou, Hao ChenICCV 2023 · 被引用 123 次
- Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkPaul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling 等NeurIPS 2023 · 被引用 120 次
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 被引用 80 次
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 被引用 1,025 次
- CAMEL: A Weakly Supervised Learning Framework for Histopathology Image SegmentationGang Xu, Zhigang Song, Zhuo Sun, Calvin Ku 等ICCV 2019 · 被引用 187 次
- HistoSegNet: Semantic Segmentation of Histological Tissue Type in Whole Slide ImagesLyndon Chan, Mahdi S. Hosseini, Corwyn Rowsell, Konstantinos N. Plataniotis 等ICCV 2019 · 被引用 131 次
- Multiple Instance Captioning: Learning Representations From Histopathology Textbooks and ArticlesJevgenij Gamper, Nasir M. RajpootCVPR 2021
相关 Paper
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang 等AAAI 2023 · 被引用 44 次
- Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational PathologyPei Liu, Luping Ji, Jiaxiang Gou, Bo Fu 等ICLR 2025
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang 等NeurIPS 2021 · 被引用 1,163 次
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionGuillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson 等CVPR 2024
- CO-PILOT: Dynamic Top-Down Point Cloud with Conditional Neighborhood Aggregation for Multi-Gigapixel Histopathology Image RepresentationRamin Nakhli, Allen W. Zhang, Ali Khajegili Mirabadi, Katherine Rich 等ICCV 2023 · 被引用 8 次
