Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning
Richard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen, Andrew D. Trister, Rahul G. Krishnan, Faisal Mahmood
Abstract
Vision Transformers (ViTs) and their multi-scale and hierarchical variations have been successful at capturing image representations but their use has been generally studied for low-resolution images (e.g. 256 × 256, 384 × 384). For gigapixel whole-slide imaging (WSI) in computational pathology, WSIs can be as large as 150000 × 150000 pixels at 20 × magnification and exhibit a hierarchical structure of visual tokens across varying resolutions: from 16 × 16 images capturing individual cells, to 4096 × 4096 images characterizing interactions within the tissue microenvironment. We introduce a new ViT architecture called the Hierarchical Image Pyramid Transformer (HIPT), which leverages the natural hierarchical structure inherent in WSIs using two levels of self-supervised learning to learn high-resolution image representations. HIPT is pretrained across 33 cancer types using 10,678 gigapixel WSIs, 408,218 4096 × 4096 images, and 104M 256 × 256 images. We benchmark HIPT representations on 9 slide-level tasks, and demonstrate that: 1) HIPT with hierarchical pretraining outperforms current state-of-the-art methods for cancer subtyping and survival prediction, 2) self-supervised ViTs are able to model important inductive biases about the hierarchical structure of phenotypes in the tumor microenvironment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a751438-812b-4e87-a7da-599fea223876Cited by top-tier papers97
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
- Cross-Modal Translation and Alignment for Survival AnalysisFengtao Zhou, Hao ChenICCV 2023 · 123 citations
- Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationWenhao Tang, Sheng Huang, Xiaoxian Zhang, Fengtao Zhou et al.ICCV 2023 · 84 citations
- The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Kexue Fu, Manning Wang et al.NeurIPS 2023 · 75 citations
- Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyWenhao Tang, Fengtao Zhou, Sheng Huang, Xiang Zhu et al.CVPR 2024 · 70 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang et al.AAAI 2023 · 44 citations
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan et al.CVPR 2026 · 1 citation
- Explainable Survival Analysis with Convolution-Involved Vision TransformerYifan Shen, Li Liu, Zhihao Tang, Zongyi Chen et al.AAAI 2022 · 27 citations
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen et al.ICCV 2021 · 369 citations
- MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in MicroscopyAlbert Dominguez Mantes, Gioele La Manno, Martin WeigertCVPR 2026 · 1 citation
