PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
K. Lokesh, Abhirama Subramanyam Penamakuri, Uday Agarwal, Apoorva Challa, Shreya K. Gowda, Somesh Gupta, Anand Mishra
Abstract
Traditionally, AI research in medical diagnosis has largely centered on image analysis. While this has led to notable advancements, the absence of patient-reported symptoms continues to hinder diagnostic accuracy. To address this, we propose a Pre-Consultation Dialogue Framework (PCDF) that mimics real-world diagnostic procedures, where doctors iteratively query patients before reaching a conclusion. Specifically, we simulate diagnostic dialogues between two vision–language models (VLMs): a DocVLM, which generates follow-up questions based on the image and dialogue history, and a PatientVLM, which responds using a symptom profile derived from the ground-truth diagnosis. We additionally conducted a small-scale clinical validation of the synthetic symptoms generated by our framework, with licensed clinicians confirming their clinical relevance, symptom coverage, and overall realism. These findings indicate that the resulting DocVLM–PatientVLM interactions form coherent, multi-turn consultations paired with images and diagnoses, which we then use to fine-tune the DocVLM. This dialogue-based supervision leads to substantial gains over image-only training, highlighting the value of realistic symptom elicitation for diagnosis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85150e9d-ccd5-468a-89c8-e7c5bdd71b74Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-World Multi-Turn DialogueSonghua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou et al.AAAI 2024 · 227 citations
Related papers
- 3MDBench: Medical Multimodal Multi-agent Dialogue BenchmarkIvan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova et al.EMNLP 2025
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa et al.CHI 2024 · 81 citations
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded ReasoningJiayuan Zhu, Jiazhen Pan, Yuyuan Liu, Fenglin Liu et al.EMNLP 2025 · 1 citation
- DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent CollaborationZhihao Jia, Mingyi Jia, Junwen Duan, Jian-xin WangEMNLP 2025 · 2 citations
- Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical NotesYang Zhou, Zhenting Sheng, Mingrui Tan, Yuting Song et al.AAAI 2026
