Bridging Radiology and Pathology Foundation Models via Concept-Based Multimodal Co-Adaptation
Yihang Chen, Yanyan Huang, Fuying Wang, Maximus Yeung, Yuming Jiang, Shujun Wang, Lequan Yu
Abstract
Pretrained medical foundation models (FMs) have shown strong generalization across diverse imaging tasks, such as disease classification in radiology and tumor grading in histopathology. While recent advances in parameter-efficient finetuning have enabled effective adaptation of FMs to downstream tasks, these approaches are typically designed for a single modality. In contrast, many clinical workflows rely on joint diagnosis from heterogeneous domains, such as radiology and pathology, where fully leveraging the representation capacity of multiple FMs remains an open challenge. To address this gap, we propose Concept Tuning and Fusing (CTF), a parameter-efficient framework that uses clinically grounded concepts as a shared semantic interface to enable cross-modal coadaptation before fusion. By incorporating task-specific concepts that are relevant across modalities, CTF aligns radiology and pathology representations, thereby enhancing their complementarity and enabling interpretation. We further design a Global-Context-Shared Prompt (GCSP) mechanism, which employs a small set of learnable tokens to capture domain-specific priors, shared patient-level information, and cross-domain context. The resulting concept alignment scores from each modality are then fused to produce a final prediction. Extensive experiments demonstrate that CTF outperforms strong unimodal, latent-fusion, and adapterbased baselines (e.g., AUC 0.903 on TCGA-GBMLGG). Notably, CTF achieves these gains without finetuning the full FMs, requiring only 0.15% additional parameters, thus highlighting the effectiveness of concept-based multimodal coadaptation. The code is available at https://github.com/HKU-MedAI/ CTF . * Corresponding author. 1 We use the term "domain" to refer to pathology and radiology, and "modal" to refer to texts and images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionYingxue Xu, Hao ChenICCV 2023 · 132 citations
Related papers
- Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed TomographyBowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang et al.ICML 2026
- Memory-Efficient Prompt Tuning for Incremental Histopathology ClassificationYu Zhu, Kang Li, Lequan Yu, Pheng-Ann HengAAAI 2024 · 4 citations
- Modaltune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-Task Learning in Digital PathologyVishwesh Ramanathan, Tony Xu, Pushpak Pati, Faruk Ahmed et al.ICCV 2025 · 4 citations
- Parameter-efficient Tuning of Large-scale Multimodal Foundation ModelHaixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin et al.NeurIPS 2023 · 45 citations
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical DiagnosisHaolin Li, Tianjie Dai, Zhe Chen, Siyuan Du et al.NeurIPS 2025 · 3 citations
