Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Mehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna, Linda G. Shapiro
摘要
Diagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology multimodal models. Training multi-model models for histopathology requires instruction tuning datasets, which currently contain information for individual image patches, without a spatial grounding of the concepts within each patch and without a wider view of the WSI. To bridge this gap, we introduce QUILT-INSTRUCT, a large-scale dataset of107, 131 histopathology-specific instruction question/answer pairs, grounded within diagnostically relevant image patches that make up the WSI. Our dataset is collected by leveraging educational histopathology videos from YouTube, which provides spatial localization of narrations by automatically extracting the narrators' cursor positions. QUILT-INSTRUCT supports contextual reasoning by extracting diagnosis and supporting facts from the entire WSI. Using QUILT-INSTRUCT, we train QUILT-LLAVA, which can reason beyond the given single image patch, enabling diagnostic reasoning across patches. To evaluate QUILT-LLAVA, we propose a compre-hensive evaluation dataset created from 985 images and 1283 human-generated question-answers. We also thor-oughly evaluate QUILT-LLAVA using public histopathology datasets, where QUILT-LLAVA significantly outperforms SOTA by over 10% on relative GPT-4 score and 4% and 9% on open and closed set VQA<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Our code, data, and model is publicly accessible at quilt-llava.github.io..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic LogicYuxuan Sun, Yixuan Si, Chenglu Zhu, Kai Zhang 等NeurIPS 2025 · 被引用 30 次
- Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert ReasonerWenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng 等AAAI 2026 · 被引用 17 次
- MedMO: Grounding and Understanding Multimodal Large Language Model for Medical ImagesAnkan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal 等CVPR 2026 · 被引用 13 次
- PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to HistopathologyFatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom Oluchi Ikezogwo 等ICCV 2025 · 被引用 13 次
- WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageYuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding 等ICCV 2025 · 被引用 10 次
它引用的顶会 Paper8
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 等NeurIPS 2021 · 被引用 463 次
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 被引用 262 次
相关 Paper
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingYing Chen, Guoan Wang, Yuanfeng Ji, Yanjun Li 等CVPR 2025
- MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image UnderstandingBasit Alawode, Arif Mahmood, Muaz Radi, Shahad Albastaki 等CVPR 2026 · 被引用 3 次
- PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationYuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu 等ICLR 2025
- CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational PathologyYuxuan Sun, Yixuan Si, Chenglu Zhu, Xuan Gong 等CVPR 2025
- LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal ModelsFeng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang 等ICLR 2025
