Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image Classification
Yingfan Ma, Xiaoyuan Luo, Kexue Fu, Manning Wang
Abstract
Pathological images play a vital role in clinical cancer diagnosis. Computer-aided diagnosis utilized on digital Whole Slide Images (WSIs) has been widely studied. The major challenge of using deep learning models for WSI analysis is the huge size of WSI images and existing methods struggle between end-to-end learning and proper modeling of contextual information. Most state-of-the-art methods utilize a two-stage strategy, in which they use a pre-trained model to extract features of small patches cut from a WSI and then input these features into a classification model. These methods can not perform end-to-end learning and consider contextual information at the same time. To solve this problem, we propose a framework that models a WSI as a pathologist's observing video and utilizes Transformer to process video clips with a divide-and-conquer strategy, which helps achieve both context-awareness and end-to-end learning. Extensive experiments on three public WSI datasets show that our proposed method outperforms existing SOTA methods in both WSI classification and positive region detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b888667e-d669-4d40-971c-b67170d1dd21Cited by top-tier papers1
Ask how each one uses itBuilds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude LearningYu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh et al.ICCV 2021 · 495 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
Related papers
- Multi-Stage Pathological Image Classification Using Semantic SegmentationShusuke Takahama, Yusuke Kurose, Yusuke Mukuta, Hiroyuki Abe et al.ICCV 2019 · 53 citations
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan et al.CVPR 2026 · 1 citation
- Explainable Survival Analysis with Convolution-Involved Vision TransformerYifan Shen, Li Liu, Zhihao Tang, Zongyi Chen et al.AAAI 2022 · 27 citations
- MulGT: Multi-Task Graph-Transformer with Task-Aware Knowledge Injection and Domain Knowledge-Driven Pooling for Whole Slide Image AnalysisWeiqin Zhao, Shujun Wang, Maximus C. F. Yeung, Tianye Niu et al.AAAI 2023 · 15 citations
- Multi-scale Domain-adversarial Multiple-instance CNN for Cancer Subtype Classification with Unannotated Histopathological ImagesNoriaki Hashimoto, Daisuke Fukushima, Ryoichi Koga, Yusuke Takagi et al.CVPR 2020
