Sparse Multi-Modal Graph Transformer with Shared-Context Processing for Representation Learning of Giga-pixel Images
Ramin Nakhli, Puria Azadi Moghadam, Haoyang Mi, Hossein Farahani, Alexander Baras, C. Blake Gilks, Ali Bashashati
Abstract
Processing giga-pixel whole slide histopathology images (WSI) is a computationally expensive task. Multiple instance learning (MIL) has become the conventional approach to process WSIs, in which these images are split into smaller patches for further processing. However, MILbased techniques ignore explicit information about the individual cells within a patch. In this paper, by defining the novel concept of shared-context processing, we designed a multi-modal Graph Transformer (AMIGO) that uses the celluar graph within the tissue to provide a single representation for a patient while taking advantage of the hierarchical structure of the tissue, enabling a dynamic focus between cell-level and tissue-level information. We benchmarked the performance of our model against multiple state-of-the-art methods in survival prediction and showed that ours can significantly outperform all of them including hierarchical Vision Transformer (ViT). More importantly, we show that our model is strongly robust to missing information to an extent that it can achieve the same performance with as low as 20% of the data. Finally, in two different cancer datasets, we demonstrated that our model was able to stratify the patients into low-risk and high-risk groups while other state-of-the-art methods failed to achieve this goal. We also publish a large dataset of immunohistochemistry images (InUIT) containing 1,600 tissue microarray (TMA) cores from 188 patients along with their survival information, making it one of the largest publicly available datasets in this context link to the request form 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2026df5c-5c47-41e5-95f1-73c2e49a8619Cited by top-tier papers4
- ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image ClassificationJiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng et al.CVPR 2024 · 38 citations
- CO-PILOT: Dynamic Top-Down Point Cloud with Conditional Neighborhood Aggregation for Multi-Gigapixel Histopathology Image RepresentationRamin Nakhli, Allen W. Zhang, Ali Khajegili Mirabadi, Katherine Rich et al.ICCV 2023 · 8 citations
- MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image ClassificationJiahao Xu, Sheng Huang, Xin Zhang, Zhixiong Nan et al.CVPR 2026 · 2 citations
- Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI AnalysisKunming Tang, Zhiguo Jiang, Jun Shi, Wei Wang et al.AAAI 2025 · 1 citation
Builds on13
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- Adversarial Graph Augmentation to Improve Graph Contrastive LearningSusheel Suresh, Pan Li, Cong Hao, Jennifer NevilleNeurIPS 2021 · 475 citations
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
- Graph Posterior Network: Bayesian Predictive Uncertainty for Node ClassificationMaximilian Stadler, Bertrand Charpentier, Simon Geisler, Daniel Zügner et al.NeurIPS 2021 · 133 citations
Related papers
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang et al.AAAI 2023 · 44 citations
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen et al.ICCV 2021 · 369 citations
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan et al.CVPR 2026 · 1 citation
- Explainable Survival Analysis with Convolution-Involved Vision TransformerYifan Shen, Li Liu, Zhihao Tang, Zongyi Chen et al.AAAI 2022 · 27 citations
- Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionYingxue Xu, Hao ChenICCV 2023 · 132 citations
