Multimodal Prototyping for cancer survival prediction
Andrew H. Song, Richard J. Chen, Guillaume Jaume, Anurag J. Vaidya, Alexander S. Baras, Faisal Mahmood
摘要
Multimodal survival methods combining gigapixel histology whole-slide images (WSIs) and transcriptomic profiles are particularly promising for patient prognostication and stratification. Current approaches involve tokenizing the WSIs into smaller patches (> 10 4 patches) and transcriptomics into gene groups, which are then integrated using a Transformer for predicting outcomes. However, this process generates many tokens, which leads to high memory requirements for computing attention and complicates post-hoc interpretability analyses. Instead, we hypothesize that we can: (1) effectively summarize the morphological content of a WSI by condensing its constituting tokens using morphological prototypes, achieving more than 300× compression; and (2) accurately characterize cellular functions by encoding the transcriptomic profile with biological pathway prototypes, all in an unsupervised fashion. The resulting multimodal tokens are then processed by a fusion network, either with a Transformer or an optimal transport cross-alignment, which now operates with a small and fixed number of tokens without approximations. Extensive evaluation on six cancer types shows that our framework outperforms state-of-the-art methods with much less computation while unlocking new interpretability analyses. The code is available at https: //github.com/mahmoodlab/MMP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution LearningChuhang Zheng, Chunwei Tian, Jie Wen, Daoqiang Zhang 等ACM MM 2025 · 被引用 13 次
- Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance LearningDaniel Shao, Joel Runevic, Richard J. Chen, Drew F. K. Williamson 等ICLR 2026 · 被引用 3 次
- Act Like a Pathologist: Tissue-Aware Whole Slide Image ReasoningWentao Huang, Weimin Lyu, Peiliang Lou, Qingqiao Hu 等CVPR 2026 · 被引用 3 次
- ConSurv: Multimodal Continual Learning for Survival AnalysisDianzhi Yu, Conghao Xiong, Yankai Chen, Wenqian Cui 等AAAI 2026 · 被引用 2 次
- Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image AnalysisXiwen Chen, Peijie Qiu, Wenhui Zhu, Hao Wang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang 等NeurIPS 2021 · 被引用 1,163 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen 等ICCV 2021 · 被引用 369 次
相关 Paper
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionGuillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson 等CVPR 2024
- PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival PredictionManahil Raza, Ayesha Azam, Talha Qaiser, Nasir M. RajpootICCV 2025 · 被引用 2 次
- Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyAndrew H. Song, Richard J. Chen, Tong Ding, Drew F. K. Williamson 等CVPR 2024 · 被引用 51 次
- Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionYingxue Xu, Hao ChenICCV 2023 · 被引用 132 次
- Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image ClassificationZhenfeng Zhuang, Fangyu Zhou, Liansheng WangAAAI 2026 · 被引用 1 次
