Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction
Hao Li, Ying Chen, Yifei Chen, Rongshan Yu, Wenxian Yang, Liansheng Wang, Bowen Ding, Yuchen Han
摘要
Whole Slide Image (WSI) classification is often formulated as a Multiple Instance Learning (MIL) problem. Recently, Vision-Language Models (VLMs) have demonstrated remarkable performance in WSI classification. However, existing methods leverage coarse-grained pathogenetic descriptions for visual representation supervision, which are insufficient to capture the complex visual appearance of pathogenetic images, hindering the generalizability of models on diverse downstream tasks. Additionally, processing high-resolution WSIs can be computationally expensive. In this paper, we propose a novel "Fine-grained Visual-Semantic Interaction" (FiVE) framework for WSI classification. It is designed to enhance the model's generalizability by leveraging the interaction between localized visual patterns and fine-grained pathological semantics. Specifically, with meticulously designed queries, we start by utilizing a large language model to extract fine-grained pathological descriptions from various non-standardized raw reports. The output descriptions are then reconstructed into fine-grained labels used for training. By introducing a Task-specific Fine-grained Semantics (TFS) module, we enable prompts to capture crucial visual information in WSIs, which enhances representation learning and augments generalization capabilities significantly. Furthermore, given that pathological visual patterns are redundantly distributed across tissue slices, we sample a subset of visual instances during training. Our method demonstrates robust generalizability and strong transferability, dominantly outperforming the counterparts on the TCGA Lung Cancer dataset with at least 9.19% higher accuracy in few-shot experiments. The code is available at: https://github.com/ls1rius/WSI FiVE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Free Lunch in Pathology Foundation Model: Task-specific Model Adaptation with Concept-Guided Feature EnhancementYanyan Huang, Weiqin Zhao, Yihang Chen, Yu Fu 等NeurIPS 2024 · 被引用 16 次
- Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image ClassificationJiaxiang Gou, Luping Ji, Pei Liu, Mao YeAAAI 2025 · 被引用 13 次
- Revisiting End-to-End Learning with Slide-level Supervision in Computational PathologyWenhao Tang, Rong Qin, Heng Fang, Fengtao Zhou 等NeurIPS 2025 · 被引用 10 次
- PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationHonglin Li, Zhongyi Shui, Yunlong Zhang, Chenglu Zhu 等NeurIPS 2025 · 被引用 6 次
- Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease ClassificationHang Guo, Qing Zhang, Zixuan Gao, Siyuan Yang 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
相关 Paper
- Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and ModelingBryan Wong, Jongwoo Kim, Huazhu Fu, Mun Yong YiNeurIPS 2025 · 被引用 4 次
- MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image ClassificationJiahao Xu, Sheng Huang, Xin Zhang, Zhixiong Nan 等CVPR 2026 · 被引用 2 次
- ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image ClassificationJiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng 等CVPR 2024 · 被引用 38 次
- MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image ClassificationJunjie Zhou, Wei Shao, Yagao Yue, Wei Mu 等NeurIPS 2025 · 被引用 1 次
- The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Kexue Fu, Manning Wang 等NeurIPS 2023 · 被引用 75 次
