Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning
Richard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen, Andrew D. Trister, Rahul G. Krishnan, Faisal Mahmood
摘要
Vision Transformers (ViTs) and their multi-scale and hierarchical variations have been successful at capturing image representations but their use has been generally studied for low-resolution images (e.g. 256 × 256, 384 × 384). For gigapixel whole-slide imaging (WSI) in computational pathology, WSIs can be as large as 150000 × 150000 pixels at 20 × magnification and exhibit a hierarchical structure of visual tokens across varying resolutions: from 16 × 16 images capturing individual cells, to 4096 × 4096 images characterizing interactions within the tissue microenvironment. We introduce a new ViT architecture called the Hierarchical Image Pyramid Transformer (HIPT), which leverages the natural hierarchical structure inherent in WSIs using two levels of self-supervised learning to learn high-resolution image representations. HIPT is pretrained across 33 cancer types using 10,678 gigapixel WSIs, 408,218 4096 × 4096 images, and 104M 256 × 256 images. We benchmark HIPT representations on 9 slide-level tasks, and demonstrate that: 1) HIPT with hierarchical pretraining outperforms current state-of-the-art methods for cancer subtyping and survival prediction, 2) self-supervised ViTs are able to model important inductive biases about the hierarchical structure of phenotypes in the tumor microenvironment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper97
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- Cross-Modal Translation and Alignment for Survival AnalysisFengtao Zhou, Hao ChenICCV 2023 · 被引用 123 次
- Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationWenhao Tang, Sheng Huang, Xiaoxian Zhang, Fengtao Zhou 等ICCV 2023 · 被引用 84 次
- The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Kexue Fu, Manning Wang 等NeurIPS 2023 · 被引用 75 次
- Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyWenhao Tang, Fengtao Zhou, Sheng Huang, Xiang Zhu 等CVPR 2024 · 被引用 70 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide ImageZhuchen Shao, Yang Chen, Hao Bian, Jian Zhang 等AAAI 2023 · 被引用 44 次
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan 等CVPR 2026 · 被引用 1 次
- Explainable Survival Analysis with Convolution-Involved Vision TransformerYifan Shen, Li Liu, Zhihao Tang, Zongyi Chen 等AAAI 2022 · 被引用 27 次
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen 等ICCV 2021 · 被引用 369 次
- MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in MicroscopyAlbert Dominguez Mantes, Gioele La Manno, Martin WeigertCVPR 2026 · 被引用 1 次
