ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification
Jiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng, Huazhu Fu
摘要
Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a substantial number of bag-level labels and solely learn from the original slides, which are easily affected by variations in data distribution. Recently, vision language model (VLM)-based methods introduced the language prior by pre-training on large-scale pathological image-text pairs. However, the previous text prompt lacks the consideration of pathological prior knowledge, there-fore does not substantially boost the model's performance. Moreover, the collection of such pairs and the pre-training process are very time-consuming and source-intensive. To solve the above problems, we propose a dual-scale vision-language multiple instance learning (ViLa-MIL) framework for whole slide image classification. Specifically, we propose a dual-scale visual descriptive text prompt based on the frozen large language model (LLM) to boost the performance of VLM effectively. To transfer the VLM to process WSI efficiently, for the image branch, we propose a prototype-guided patch decoder to aggregate the patch features progressively by grouping similar patches into the same prototype; for the text branch, we introduce a context-guided text decoder to enhance the text features by incorporating the multi-granular image contexts. Extensive studies on three multi-cancer and multi-center subtyping datasets demonstrate the superiority of ViLa-MIL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image AnalysisDi Zhang, Zhangpeng Gong, Xiaobo Pang, Jiashuai Liu 等CVPR 2026 · 被引用 11 次
- Revisiting End-to-End Learning with Slide-level Supervision in Computational PathologyWenhao Tang, Rong Qin, Heng Fang, Fengtao Zhou 等NeurIPS 2025 · 被引用 10 次
- DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival PredictionYucheng Xing, ling huang, Jingying Ma, Ruping Hong 等ICML 2026 · 被引用 8 次
- PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationHonglin Li, Zhongyi Shui, Yunlong Zhang, Chenglu Zhu 等NeurIPS 2025 · 被引用 6 次
- Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and ModelingBryan Wong, Jongwoo Kim, Huazhu Fu, Mun Yong YiNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
相关 Paper
- MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image ClassificationJunjie Zhou, Wei Shao, Yagao Yue, Wei Mu 等NeurIPS 2025 · 被引用 1 次
- Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image ClassificationJiaxiang Gou, Luping Ji, Pei Liu, Mao YeAAAI 2025 · 被引用 13 次
- Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational PathologyPei Liu, Luping Ji, Jiaxiang Gou, Bo Fu 等ICLR 2025
- Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic InteractionHao Li, Ying Chen, Yifei Chen, Rongshan Yu 等CVPR 2024
- Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image ClassificationZhenfeng Zhuang, Fangyu Zhou, Liansheng WangAAAI 2026 · 被引用 1 次
