OmniDoctor: Towards LLM-centric Lifelong Learning for New Emerging Medical VQA Tasks
Na Jiang, Wenhui Zheng, Xuqian Gu, Jingjing Wang
Abstract
Current medical multimodal large language models (MLLMs) have demonstrated high accuracy and effectiveness on specific medical visual question answering (Medical VQA) tasks. However, they largely fail to tackle continuously emerging unseen Medical VQA scenarios (e.g., MRI, X-ray) in real-world settings, which significantly hinders their broader adoption in practical clinical environments. Motivated by these gaps, this paper introduces a new task, namely LLM-centric Lifelong Learning for New Medical VQA (L3NMV), which enables Large Language Models (LLMs) to continually learn medical image-text knowledge across various medical VQA tasks. Furthermore, this paper reveals two critical challenges: 1) Efficient medical knowledge retention (Each-task), which aims to retain essential knowledge for each Medical VQA task efficiently with limited data. 2) Efficient medical interference mitigation (Cross-task), which focuses on efficiently mitigating information interference across various Medical VQA tasks with knowledge barriers. To address these challenges, this paper proposes the OmniDoctor model, i.e., an omniscient doctor that simulates how doctors continuously update their knowledge and skills through continuous medical education, with the goal of equipping the model with lifelong learning capabilities via an efficient incremental medical parameter constraining mechanism for L3 NMV. This model is designed with two key modules to address the above two challenges, respectively. Especially, this paper constructs an Unseen L3NMV dataset to simulate real-world incremental clinical scenarios. Extensive experiments on this dataset demonstrate that OmniDoctor outperforms several advanced lifelong learning baselines. These results justify the significance of the L3 NMV task and the effectiveness of OmniDoctor in continually adapting to new Medical VQA tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao et al.CVPR 2024
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical TasksZhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang et al.CVPR 2026 · 7 citations
- Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoEXun Zhu, Ying Hu, Fanbin Mo, Miao Li et al.NeurIPS 2024 · 29 citations
- Sparse Spectral LoRA: Routed Experts for Medical VLMsOmid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao et al.CVPR 2026 · 3 citations
- MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language ModelsDexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao et al.AAAI 2026 · 1 citation
