Lune

NeurIPS2025Top-tier venue

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

David Dai, Peilin Chen, Chanakya Ekbote, Paul Pu Liang

2025Year
48Citations
9Top-tier citations

Abstract

Clinical decision-making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision-centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time-series signals, and text reports. QoQ-Med is trained with Domain-aware Relative Policy Optimization (DRPO), a novel reinforcement-learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro-F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and down- stream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces at this link. necessary component for responsible clinical implementation, regulatory compliance, and effective human-AI collaboration in healthcare environments [6,66,11].

In this work, we introduce QoQ-Med: a generalist clinical multimodal foundation model with precise reasoning capabilities spanning clinical images, time series data, and textual records across 9 clinical domains. Our work makes two primary contributions:

  1. Firstly, to tackle the challenges associated with balancing heterogeneous data for balanced and efficient training across 1D to 3D data, we propose Domain-aware Group Relative Policy Optimization (DRPO). DRPO employs hierarchical scaling based on the domain of the input data, which encourages the model's learning on scarce and hard domains, allowing balanced learning across difficulty levels. Our empirical evaluation demonstrates that DRPO consistently outperforms established RL approaches in diverse multi-domain settings, with up to 43% improvement in average F1 score across 8 clinical vision modalities. 2. To tackle the second challenge of expert interpretability, we design and release one of the first multimodal clinical reasoning models, namely QoQ-Med-7B/32B (Qwen Omni-Reasoning on Medical Questions), that integrates visual, time series, and textual data for comprehensive analysis of clinical records, facilitating more holistic diagnostic reasoning. QoQ-Med is trained to highlight salient regions in the visual input data, advancing the interpretability while allowing the clinician to check the model's diagnosis with ease. To the best of our knowledge, QoQ-Med is currently the largest open-source multimodal reasoning model for clinical diagnosis, and the only MLLM that integrates time series data (ECG) with traditional clinical vision modalities.

Finally, we publicly release our model, training pipeline, and reasoning traces generated by the model across 2.61 million question-answer pairs at this link. This marks one of the largest resources for transparent and reproducible multimodal reasoning in the clinical domain.

2 Related Work

Recent work has adapted vision-language interfaces to the medical domain, yielding models such as LLaVa-Med [48], RadLM [90],. These models couple frozen LLM backbones with image encoders and are trained on radiology or pathology visual-question-answering and reportgeneration benchmarks [24,39,92,88,86]. Although these systems demonstrate impressive zeroshot understanding, their training corpora are dominated by single-institution chest X-rays, retinal photographs, and pathology slides, resulting in limited generalization to demographic diversity and poor robustness to real-world distribution [44,63,51]. GEM [47] is the only MLLM incorporating ECG data, but the training focus is purely ECG, which does not provide a comprehensive diagnosis aggregating multiple sources. Our work addresses these gaps by assembling a richer corpus spanning imaging, time-series, and text, and by designing an architecture that natively models medical timeseries alongside traditional modalities.

The introduction of instruction tuning precipitated a rapid shift from supervised fine-tuning to reinforcement learning pipelines. Proximal Policy Optimization (PPO) [74] as popularized by InstructGPT, trains LLMs against a reward model under a KL penalty to a frozen reference, with an auxiliary critic estimating advantages [62]. While effective, PPO's critic incurs substantial memory and computation costs and can destabilise multi-task optimization [73]. To reduce overhead, critic-free objectives such as Direct Preference Optimization (DPO) [70] and Group Relative Policy Op

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext edb27f2c-19c9-4223-8311-edcd8dbf6dd4

Cited by top-tier papers9

Ask how each one uses it

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines