Multi-modal Medical Diagnosis via Large-small Model Collaboration
Wanyi Chen, Zihua Zhao, Jiangchao Yao, Ya Zhang, Jiajun Bu, Haishuai Wang
Abstract
Recent advances in medical AI have shown a clear trend towards large models in healthcare. However, developing large models for multi-modal medical diagnosis remains challenging due to a lack of sufficient modal-complete medical data. Most existing multi-modal diagnostic models are relatively small and struggle with limited feature extraction capabilities. To bridge this gap, we propose AdaCoMed, an adaptive collaborative-learning framework that synergistically integrates the off-the-shelf medical single-modal large models with multi-modal small models. Our framework first employs a mixture-of-modality-experts (MoME) architecture to combine features extracted from multiple single-modal medical large models, and then introduces a novel adaptive co-learning mechanism to collaborate with a multi-modal small model. This co-learning mechanism, guided by an adaptive weighting strategy, dynamically balances the complementary strengths between the MoMEfused large model features and the cross-modal reasoning capabilities of the small model. Extensive experiments on two representative multi-modal medical datasets (MIMIC-IV-MM and MMIST ccRCC) across six modalities and four diagnostic tasks demonstrate consistent improvements over state-of-the-art baselines, making it a promising solution for real-world medical diagnosis applications. The code is available at https://github.com/Zoew420/ AdaCoMed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb5fce4d-8f4e-4964-96bf-b806b8861fb0Cited by top-tier papers2
- Differential-Informed Sample Selection Accelerates Multimodal Contrastive LearningZihua Zhao, Feng Hong, Mengxi Chen, Pengyi Chen et al.ICCV 2025
- Unifying Multi-View Knowledge for Graph Learning via Model CollaborationZhihao Wu, Jielong Lu, Zihan Fang, Jinyu Cai et al.AAAI 2026
Builds on14
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
- CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionJie Liu, Yixiao Zhang, Jieneng Chen, Junfei Xiao et al.ICCV 2023 · 336 citations
- Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasZhongwei Wan, Che Liu, Mi Zhang, Jie Fu et al.NeurIPS 2023 · 114 citations
- A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesMingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang et al.AAAI 2024 · 71 citations
Related papers
- Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-ExpertsSukwon Yun, Inyoung Choi, Jie Peng, Yangfan Wu et al.NeurIPS 2024 · 98 citations
- A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingXiao Wu, Ting-Zhu Huang, Liang-Jian Deng, Yanyuan Qiao et al.EMNLP 2025 · 1 citation
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- Dynamic Modeling of Patients, Modalities and Tasks via Multi-modal Multi-task Mixture of ExpertsChenwei Wu, Zitao Shuai, Zhengxu Tang, Luning Wang et al.ICLR 2025
- Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisBing Cao, Han Zhang, Nannan Wang, Xinbo Gao et al.AAAI 2020 · 94 citations
