FBTA: Enabling Single-GPU End-to-End Gigapixel WSI Classification with Feature Bridging and Translation Alignment
Jiuyang Dong, Jiahan Li, Junjun Jiang, Yongbing Zhang
Abstract
Whole-slide images (WSIs) in computational pathology contain billions of pixels, making end-to-end training of feature extractors and multi-instance learning (MIL) networks infeasible on a single commodity GPU.Existing methods often freeze the feature extractor and train MIL networks on the resulting frozen features, which introduces a semantic gap that limits downstream performance. To address this issue, we propose FBTA, a Feature Bridging and Translation Alignment framework for WSI classification. FBTA is the first end-to-end MIL framework trainable on a single 24 GB GPU, leveraging three complementary feature-bag views: end-to-end features enable joint optimization, frozen features stabilize training, and translated features support practical inference.Experiments on diverse datasets, including TCGA-NSCLC (Shot20/50/100) and TCGA-STAD, demonstrate the effectiveness and generality of FBTA, which consistently improves performance across three MIL architectures and two extractors. For example, with ResNet-50 as the extractor, FBTA improves the accuracy of the classic ABMIL by 13.1% and 15.8% on the NSCLC-Shot50 and TCGA-STAD datasets, respectively, and further enhances the state-of-the-art MambaMIL by 4.1% and 9.2% on the same datasets. Moreover, FBTA yields additional gains for MIL models that incorporate self-supervised pretraining strategies and data augmentation techniques.These results suggest FBTA is a feasible and scalable framework for end-to-end MIL on gigapixel WSIs. The code will be available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b5c2e00-5d26-4ba9-8cff-9b2991da7129Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
Related papers
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan et al.CVPR 2026 · 1 citation
- Revisiting End-to-End Learning with Slide-level Supervision in Computational PathologyWenhao Tang, Rong Qin, Heng Fang, Fengtao Zhou et al.NeurIPS 2025 · 10 citations
- Do Multiple Instance Learning Models Transfer?Daniel Shao, Richard J. Chen, Andrew H. Song, Joel Runevic et al.ICML 2025
- SAM-MIL: A Spatial Contextual Aware Multiple Instance Learning Approach for Whole Slide Image ClassificationHeng Fang, Sheng Huang, Wenhao Tang, Luwen Huangfu et al.ACM MM 2024 · 13 citations
- ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image ClassificationJiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng et al.CVPR 2024 · 38 citations
