Explainable Survival Analysis with Convolution-Involved Vision Transformer
Yifan Shen, Li Liu, Zhihao Tang, Zongyi Chen, Guixiang Ma, Jiyan Dong, Xi Zhang, Lin Yang, Qingfeng Zheng
Abstract
Image-based survival prediction models can facilitate doctors in diagnosing and treating cancer patients. With the advance of digital pathology technologies, the big whole slide images (WSIs) provide increased resolution and more details for diagnosis. However, the gigabytesize or even terabyte-size WSIs would make most models computationally infeasible. To this end, instead of using the complete WSIs, most of the existing models only use a pre-selected subset of key patches or patch clusters as input, which might discard some important morphology information. In this work, we propose a novel survival analysis model to fully utilize the complete WSI information. We show that the use of a Vision Transformer (ViT) backbone, together with convolution operations involved in it, is an effective approach to improve the prediction performance. Additionally, we present a post-hoc explainable method to identify the most salient patches and distinct morphology features, making the model more faithful and the results easier to comprehend by human users. Evaluations on two large cancer datasets show that our proposed model is more effective and has better interpretability for survival prediction. We would make the code publicly available upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b77ddc2-04d6-4cf8-a88b-d1aa208f6753Cited by top-tier papers3
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang et al.AAAI 2023 · 44 citations
- Point Transformer with Federated Learning for Predicting Breast Cancer HER2 Status from Hematoxylin and Eosin-Stained Whole Slide ImagesBao Li, Zhenyu Liu, Lizhi Shao, Bensheng Qiu et al.AAAI 2024 · 13 citations
- From Representation Space to Prognostic Insights: Whole Slide Image Generation with Hierarchical Diffusion Model for Survival PredictionZhihao Tang, Xi Zhang, Chaozhuo LiAAAI 2025 · 1 citation
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Quantifying Explainers of Graph Neural Networks in Computational PathologyGuillaume Jaume, Pushpak Pati, Behzad Bozorgtabar, Antonio Foncubierta et al.CVPR 2021
Related papers
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan et al.CVPR 2026 · 1 citation
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionGuillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson et al.CVPR 2024
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationYingfan Ma, Xiaoyuan Luo, Kexue Fu, Manning WangAAAI 2024 · 10 citations
- Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational PathologyPei Liu, Luping Ji, Jiaxiang Gou, Bo Fu et al.ICLR 2025
