WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Jipeng Zhang, Xiangjian He, Song Wu, Xiaohan Xing, Sen Yang, Xiyue Wang, Linlin Shen
Abstract
Recent advances in computational pathology have introduced whole slide image (WSI)-level multimodal large language models (MLLMs) for automated pathological analysis. However, current WSI-level MLLMs face two critical challenges: limited explainability in their decision-making process and insufficient attention to morphological features crucial for accurate diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphologyaware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, specifically designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. To the best of our knowledge, WSI-Bench presents the first benchmarking systematically evaluate morphological understanding capabilities in WSI analysis. To enhance the model explainability, we present WSI-LLaVA, an MLLM framework for gigapixel WSI understanding with a three-stage training strategy, which can provide detailed morphological findings to explain its final answer. For more precise model assessment in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance, focusing on clinical accuracy. Extensive evaluation on WSI-Bench reveals both the capabilities and limitations of current WSI MLLMs in morphological analysis and various pathology tasks, while demonstrating WSI-LLaVA's superior performance across all capabilities on both internal and external datasets. Source code and data are released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c974470d-0ba8-4f0d-96ba-833b3806780bCited by top-tier papers4
- OralGPT-Omni: A Versatile Dental Multimodal Large Language ModelJing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan et al.CVPR 2026 · 11 citations
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype ControlMinghao Han, Yichen Liu, Yizhou Liu, Zizhi Chen et al.CVPR 2026 · 5 citations
- MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image UnderstandingBasit Alawode, Arif Mahmood, Muaz Radi, Shahad Albastaki et al.CVPR 2026 · 3 citations
- Act Like a Pathologist: Tissue-Aware Whole Slide Image ReasoningWentao Huang, Weimin Lyu, Peiliang Lou, Qingqiao Hu et al.CVPR 2026 · 3 citations
Builds on2
- Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology VideosMehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna et al.CVPR 2024 · 37 citations
- EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language ModelMeidan Ding, Jipeng Zhang, Wenxuan Wang, Haiqin Zhong et al.ACL 2025
Related papers
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingYing Chen, Guoan Wang, Yuanfeng Ji, Yanjun Li et al.CVPR 2025
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen et al.ACL 2026
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical TasksZhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang et al.CVPR 2026 · 7 citations
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisBo Liu, Ke Zou, Li-Ming Zhan, Zexin Lu et al.ICCV 2025 · 10 citations
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningHaozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu et al.CVPR 2026 · 15 citations
