Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification
Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang
Abstract
While Large Language Models (LLMs) are emerging as a promising direction in computational pathology, the substantial computational cost of giga-pixel Whole Slide Images (WSIs) necessitates the use of Multi-Instance Learning (MIL) to enable effective modeling. A key challenge is that pathological tasks typically provide only bag-level labels, while instance-level descriptions generated by LLMs often suffer from bias due to a lack of fine-grained medical knowledge. To address this, we propose that constructing task-specific pathological entity prototypes is crucial for learning generalizable features and enhancing model interpretability. Furthermore, existing vision-language MIL methods often employ unidirectional guidance, limiting cross-modal synergy. In this paper, we introduce a novel approach, Multimodal Prototype-based Multi-Instance Learning, that promotes bidirectional interaction through a balanced information compression scheme. Specifically, we leverage a frozen LLM to generate task-specific pathological entity descriptions, which are learned as text prototypes. Concurrently, the vision branch learns instance-level prototypes to mitigate the model's reliance on redundant data. For the fusion stage, we employ the Stereoscopic Optimal Transport (SOT) algorithm, which is based on a similarity metric, thereby facilitating broader semantic alignment in a higher-dimensional space. We conduct few-shot classification and explainability experiments on three distinct cancer datasets, and the results demonstrate the superior generalization capabilities of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61ac91ce-e6c0-4a67-806e-6a9fd989bab3Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationHongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao et al.CVPR 2022 · 402 citations
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen et al.ICCV 2021 · 369 citations
- Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionYingxue Xu, Hao ChenICCV 2023 · 132 citations
- H^2-MIL: Exploring Hierarchical Representation with Heterogeneous Multiple Instance Learning for Whole Slide Image AnalysisWentai Hou, Lequan Yu, Chengxuan Lin, Helong Huang et al.AAAI 2022 · 106 citations
Related papers
- ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image ClassificationJiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng et al.CVPR 2024 · 38 citations
- Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic InteractionHao Li, Ying Chen, Yifei Chen, Rongshan Yu et al.CVPR 2024
- MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image ClassificationJunjie Zhou, Wei Shao, Yagao Yue, Wei Mu et al.NeurIPS 2025 · 1 citation
- Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology ImagesMing Y. Lu, Bowen Chen, Andrew Zhang, Drew F. K. Williamson et al.CVPR 2023
- MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image UnderstandingBasit Alawode, Arif Mahmood, Muaz Radi, Shahad Albastaki et al.CVPR 2026 · 3 citations
