QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
David Dai, Peilin Chen, Chanakya Ekbote, Paul Pu Liang
摘要
Clinical decision-making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision-centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time-series signals, and text reports. QoQ-Med is trained with Domain-aware Relative Policy Optimization (DRPO), a novel reinforcement-learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro-F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and down- stream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces at this link. necessary component for responsible clinical implementation, regulatory compliance, and effective human-AI collaboration in healthcare environments [6,66,11].
In this work, we introduce QoQ-Med: a generalist clinical multimodal foundation model with precise reasoning capabilities spanning clinical images, time series data, and textual records across 9 clinical domains. Our work makes two primary contributions:
- Firstly, to tackle the challenges associated with balancing heterogeneous data for balanced and efficient training across 1D to 3D data, we propose Domain-aware Group Relative Policy Optimization (DRPO). DRPO employs hierarchical scaling based on the domain of the input data, which encourages the model's learning on scarce and hard domains, allowing balanced learning across difficulty levels. Our empirical evaluation demonstrates that DRPO consistently outperforms established RL approaches in diverse multi-domain settings, with up to 43% improvement in average F1 score across 8 clinical vision modalities. 2. To tackle the second challenge of expert interpretability, we design and release one of the first multimodal clinical reasoning models, namely QoQ-Med-7B/32B (Qwen Omni-Reasoning on Medical Questions), that integrates visual, time series, and textual data for comprehensive analysis of clinical records, facilitating more holistic diagnostic reasoning. QoQ-Med is trained to highlight salient regions in the visual input data, advancing the interpretability while allowing the clinician to check the model's diagnosis with ease. To the best of our knowledge, QoQ-Med is currently the largest open-source multimodal reasoning model for clinical diagnosis, and the only MLLM that integrates time series data (ECG) with traditional clinical vision modalities.
Finally, we publicly release our model, training pipeline, and reasoning traces generated by the model across 2.61 million question-answer pairs at this link. This marks one of the largest resources for transparent and reproducible multimodal reasoning in the clinical domain.
2 Related Work
Recent work has adapted vision-language interfaces to the medical domain, yielding models such as LLaVa-Med [48], RadLM [90],. These models couple frozen LLM backbones with image encoders and are trained on radiology or pathology visual-question-answering and reportgeneration benchmarks [24,39,92,88,86]. Although these systems demonstrate impressive zeroshot understanding, their training corpora are dominated by single-institution chest X-rays, retinal photographs, and pathology slides, resulting in limited generalization to demographic diversity and poor robustness to real-world distribution [44,63,51]. GEM [47] is the only MLLM incorporating ECG data, but the training focus is purely ECG, which does not provide a comprehensive diagnosis aggregating multiple sources. Our work addresses these gaps by assembling a richer corpus spanning imaging, time-series, and text, and by designing an architecture that natively models medical timeseries alongside traditional modalities.
The introduction of instruction tuning precipitated a rapid shift from supervised fine-tuning to reinforcement learning pipelines. Proximal Policy Optimization (PPO) [74] as popularized by InstructGPT, trains LLMs against a reward model under a KL penalty to a frozen reference, with an auxiliary critic estimating advantages [62]. While effective, PPO's critic incurs substantial memory and computation costs and can destabilise multi-task optimization [73]. To reduce overhead, critic-free objectives such as Direct Preference Optimization (DPO) [70] and Group Relative Policy Op
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical ReasoningTimothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin 等CVPR 2026 · 被引用 10 次
- MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement LearningZheng Jiang, Heng Guo, Chengyu Fang, Changchen Xiao 等ICLR 2026 · 被引用 8 次
- PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated ReportingDanyal Maqbool, Changhee Lee, Zachary Huemann, Samuel Church 等CVPR 2026 · 被引用 2 次
- Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report GenerationYuan Wang, Shujian Gao, Jiaxiang Liu, Songtao Jiang 等AAAI 2026 · 被引用 2 次
- Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC VideosShreshth Saini, Bowen Chen, Yilin Wang, Neil Birkbeck 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Brain Network TransformerXuan Kan, Wei Dai, Hejie Cui, Zilong Zhang 等NeurIPS 2022 · 被引用 272 次
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
相关 Paper
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing 等AAAI 2026 · 被引用 4 次
- MedGR2: Breaking the Data Barrier for Medical Reasoning via Generative Reward LearningWeihai Zhi, Jiayan Guo, Shangyang LiAAAI 2026 · 被引用 5 次
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video UnderstandingYuhao Su, Anwesa Choudhuri, Zhongpai Gao, Benjamin Planche 等CVPR 2026 · 被引用 10 次
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical ReasoningPeng Xia, Jinglu Wang, Yibo Peng, Kaide Zeng 等ICLR 2026 · 被引用 47 次
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang 等NeurIPS 2025 · 被引用 255 次
