MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
Basit Alawode, Arif Mahmood, Muaz Radi, Shahad Albastaki, Asim Khan, Muhammad Bilal, Moshira Ali Abdalla, Mohammed Bennamoun, Sajid Javed
摘要
Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic cues arise from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language Models (MLLMs) typically compress an entire WSI into a single embedding, which hinders fine-grained grounding and ignores how pathologists synthesize evidence across different scales. We introduce MLLM-HWSI, a Hierarchical WSI-level MLLM that aligns visual features with pathology language at four distinct scales—cell as word, patch as phrase, region as sentence, and WSI as paragraph—to support interpretable, evidence-grounded reasoning. MLLM-HWSI decomposes each WSI into multi-scale embeddings with scale-specific VL projectors and jointly enforces (i) a hierarchical contrastive objective and (ii) a cross-scale consistency loss, preserving semantic coherence from cells to the WSI. To make gigapixel processing tractable and clinically meaningful, we compute diagnostically relevant tokens and aggregate segmented cell embeddings into a compact cellular token per-patch using a lightweight Cell–Cell Attention Fusion (CCAF) transformer. The projected multi-scale tokens are fused with text tokens and fed to an instruction-tuned LLM for open-ended reasoning, VQA, report, and caption generation tasks. Trained in three stages, MLLM-HWSI achieves new SOTA results on 13 WSI-level benchmarks across six CPath tasks. By grounding language in calibrated, multi-scale visual evidence, HMLLM provides accurate, interpretable outputs that mirror expert diagnostic workflows and advance holistic WSI understanding. Code will be released upon the publication.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen 等CVPR 2022 · 被引用 490 次
- PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyYuxuan Sun, Chenglu Zhu, Sunyi Zheng, Kai Zhang 等AAAI 2024 · 被引用 92 次
- The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Kexue Fu, Manning Wang 等NeurIPS 2023 · 被引用 75 次
- Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology VideosMehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna 等CVPR 2024 · 被引用 37 次
- WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageYuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding 等ICCV 2025 · 被引用 10 次
相关 Paper
- PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational PathologyFengchun Liu, Songhan Jiang, Linghan Cai, Ziyue Wang 等AAAI 2026 · 被引用 2 次
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingYing Chen, Guoan Wang, Yuanfeng Ji, Yanjun Li 等CVPR 2025
- CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational PathologyYuxuan Sun, Yixuan Si, Chenglu Zhu, Xuan Gong 等CVPR 2025
- Cello: A Universal Cell-wise Feature Aggregation framework for Reliable Pathology Images AnalysisHengrui Lou, Weihan Li, Jiazhen Yang, Lingxiang Jia 等ICML 2026
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic LogicYuxuan Sun, Yixuan Si, Chenglu Zhu, Kai Zhang 等NeurIPS 2025 · 被引用 30 次
