QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
David Dai, Peilin Chen, Chanakya Ekbote, Paul Pu Liang
Abstract
Clinical decision-making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision-centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time-series signals, and text reports. QoQ-Med is trained with Domain-aware Relative Policy Optimization (DRPO), a novel reinforcement-learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro-F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and down- stream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces at this link. necessary component for responsible clinical implementation, regulatory compliance, and effective human-AI collaboration in healthcare environments [6,66,11].
In this work, we introduce QoQ-Med: a generalist clinical multimodal foundation model with precise reasoning capabilities spanning clinical images, time series data, and textual records across 9 clinical domains. Our work makes two primary contributions:
- Firstly, to tackle the challenges associated with balancing heterogeneous data for balanced and efficient training across 1D to 3D data, we propose Domain-aware Group Relative Policy Optimization (DRPO). DRPO employs hierarchical scaling based on the domain of the input data, which encourages the model's learning on scarce and hard domains, allowing balanced learning across difficulty levels. Our empirical evaluation demonstrates that DRPO consistently outperforms established RL approaches in diverse multi-domain settings, with up to 43% improvement in average F1 score across 8 clinical vision modalities. 2. To tackle the second challenge of expert interpretability, we design and release one of the first multimodal clinical reasoning models, namely QoQ-Med-7B/32B (Qwen Omni-Reasoning on Medical Questions), that integrates visual, time series, and textual data for comprehensive analysis of clinical records, facilitating more holistic diagnostic reasoning. QoQ-Med is trained to highlight salient regions in the visual input data, advancing the interpretability while allowing the clinician to check the model's diagnosis with ease. To the best of our knowledge, QoQ-Med is currently the largest open-source multimodal reasoning model for clinical diagnosis, and the only MLLM that integrates time series data (ECG) with traditional clinical vision modalities.
Finally, we publicly release our model, training pipeline, and reasoning traces generated by the model across 2.61 million question-answer pairs at this link. This marks one of the largest resources for transparent and reproducible multimodal reasoning in the clinical domain.
2 Related Work
Recent work has adapted vision-language interfaces to the medical domain, yielding models such as LLaVa-Med [48], RadLM [90],. These models couple frozen LLM backbones with image encoders and are trained on radiology or pathology visual-question-answering and reportgeneration benchmarks [24,39,92,88,86]. Although these systems demonstrate impressive zeroshot understanding, their training corpora are dominated by single-institution chest X-rays, retinal photographs, and pathology slides, resulting in limited generalization to demographic diversity and poor robustness to real-world distribution [44,63,51]. GEM [47] is the only MLLM incorporating ECG data, but the training focus is purely ECG, which does not provide a comprehensive diagnosis aggregating multiple sources. Our work addresses these gaps by assembling a richer corpus spanning imaging, time-series, and text, and by designing an architecture that natively models medical timeseries alongside traditional modalities.
The introduction of instruction tuning precipitated a rapid shift from supervised fine-tuning to reinforcement learning pipelines. Proximal Policy Optimization (PPO) [74] as popularized by InstructGPT, trains LLMs against a reward model under a KL penalty to a frozen reference, with an auxiliary critic estimating advantages [62]. While effective, PPO's critic incurs substantial memory and computation costs and can destabilise multi-task optimization [73]. To reduce overhead, critic-free objectives such as Direct Preference Optimization (DPO) [70] and Group Relative Policy Op
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edb27f2c-19c9-4223-8311-edcd8dbf6dd4Cited by top-tier papers9
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical ReasoningTimothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin et al.CVPR 2026 · 10 citations
- MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement LearningZheng Jiang, Heng Guo, Chengyu Fang, Changchen Xiao et al.ICLR 2026 · 8 citations
- PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated ReportingDanyal Maqbool, Changhee Lee, Zachary Huemann, Samuel Church et al.CVPR 2026 · 2 citations
- Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report GenerationYuan Wang, Shujian Gao, Jiaxiang Liu, Songtao Jiang et al.AAAI 2026 · 2 citations
- Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC VideosShreshth Saini, Bowen Chen, Yilin Wang, Neil Birkbeck et al.CVPR 2026 · 1 citation
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Brain Network TransformerXuan Kan, Wei Dai, Hejie Cui, Zilong Zhang et al.NeurIPS 2022 · 272 citations
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
Related papers
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- MedGR2: Breaking the Data Barrier for Medical Reasoning via Generative Reward LearningWeihai Zhi, Jiayan Guo, Shangyang LiAAAI 2026 · 5 citations
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video UnderstandingYuhao Su, Anwesa Choudhuri, Zhongpai Gao, Benjamin Planche et al.CVPR 2026 · 10 citations
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningPeng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng et al.ICLR 2026 · 47 citations
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang et al.NeurIPS 2025 · 255 citations
