Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Mehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna, Linda G. Shapiro
Abstract
Diagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology multimodal models. Training multi-model models for histopathology requires instruction tuning datasets, which currently contain information for individual image patches, without a spatial grounding of the concepts within each patch and without a wider view of the WSI. To bridge this gap, we introduce QUILT-INSTRUCT, a large-scale dataset of107, 131 histopathology-specific instruction question/answer pairs, grounded within diagnostically relevant image patches that make up the WSI. Our dataset is collected by leveraging educational histopathology videos from YouTube, which provides spatial localization of narrations by automatically extracting the narrators' cursor positions. QUILT-INSTRUCT supports contextual reasoning by extracting diagnosis and supporting facts from the entire WSI. Using QUILT-INSTRUCT, we train QUILT-LLAVA, which can reason beyond the given single image patch, enabling diagnostic reasoning across patches. To evaluate QUILT-LLAVA, we propose a compre-hensive evaluation dataset created from 985 images and 1283 human-generated question-answers. We also thor-oughly evaluate QUILT-LLAVA using public histopathology datasets, where QUILT-LLAVA significantly outperforms SOTA by over 10% on relative GPT-4 score and 4% and 9% on open and closed set VQA<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Our code, data, and model is publicly accessible at quilt-llava.github.io..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e865841b-b79e-48be-a444-0aae08e7b92aCited by top-tier papers22
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic LogicYuxuan Sun, Yixuan Si, Chenglu Zhu, Kai Zhang et al.NeurIPS 2025 · 30 citations
- Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert ReasonerWenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng et al.AAAI 2026 · 17 citations
- MedMO: Grounding and Understanding Multimodal Large Language Model for Medical ImagesAnkan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal et al.CVPR 2026 · 13 citations
- PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to HistopathologyFatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom Oluchi Ikezogwo et al.ICCV 2025 · 13 citations
- WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageYuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding et al.ICCV 2025 · 10 citations
Builds on8
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu et al.NeurIPS 2021 · 463 citations
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
Related papers
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingYing Chen, Guoan Wang, Yuanfeng Ji, Yanjun Li et al.CVPR 2025
- MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image UnderstandingBasit Alawode, Arif Mahmood, Muaz Radi, Shahad Albastaki et al.CVPR 2026 · 3 citations
- PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationYuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu et al.ICLR 2025
- CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational PathologyYuxuan Sun, Yixuan Si, Chenglu Zhu, Xuan Gong et al.CVPR 2025
- LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal ModelsFeng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang et al.ICLR 2025
