Multi-modal Vision Pre-training for Medical Image Analysis
Shaohao Rui, Lingzhi Chen, Zhenyu Tang, Lilong Wang, Mianxin Liu, Shaoting Zhang, Xiaosong Wang
摘要
Self-supervised learning has greatly facilitated medical image analysis by suppressing the training data requirement for real-world applications. Current paradigms predominantly rely on self-supervision within uni-modal image data, thereby neglecting the inter-modal correlations essential for effective learning of cross-modal image representations. This limitation is particularly significant for naturally grouped multi-modal data, e.g., multi-parametric MRI scans for a patient undergoing various functional imaging protocols in the same study. To bridge this gap, we conduct a novel multi-modal image pre-training with three proxy tasks to facilitate the learning of cross-modality representations and correlations using multi-modal brain MRI scans (over 2.4 million images in 16,022 scans of 3,755 patients), i.e., cross-modal image reconstruction, modalityaware contrastive learning, and modality template distillation. To demonstrate the generalizability of our pre-trained model, we conduct extensive experiments on various benchmarks with ten downstream tasks. The superior performance of our method is reported in comparison to state-ofthe-art pre-training methods, with Dice Score improvement of 0.28%-14.47% across six segmentation benchmarks and a consistent accuracy boost of 0.65%-18.07% in four individual image classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SegMoTE: Token-Level Mixture of Experts for Medical Image SegmentationYujie Lu, Jingwen Li, Sibo Ju, Yanzhou Su 等CVPR 2026 · 被引用 2 次
- Does YOLO Really Need to See Every Training Image in Every Epoch?Xingxing Xie, Jiahua Dong, Junwei Han, Gong ChengCVPR 2026 · 被引用 1 次
- InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-trainingZihao Luo, Shaohao Rui, Zhenyu Tang, Guotai Wang 等CVPR 2026
- Beyond the Static-World: Lifelong Learning for All-in-One Medical Image RestorationShihao Shan, Hongying Liu, Fanhua Shang, Liang Wan 等CVPR 2026
- Masked-Diffusion Autoencoders for 3D Medical Vision Representation LearningJiachen Tu, Guanghui Qin, Theodore Zhengde Zhao, Jeya Maria Jose Valanarasu 等CVPR 2026
它引用的顶会 Paper23
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth 等CVPR 2022 · 被引用 736 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
相关 Paper
- Continual Self-Supervised Learning: Towards Universal Multi-Modal Medical Data Representation LearningYiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen 等CVPR 2024 · 被引用 30 次
- Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labelsJizong Peng, Ping Wang, Christian Desrosiers, Marco PedersoliNeurIPS 2021 · 被引用 80 次
- Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical ImagingTan Pan, Shuhao Mei, Yixuan Sun, Kaiyu Guo 等ICML 2026
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu 等CVPR 2026
- LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph MatchingDuy M. H. Nguyen, Hoang Nguyen, Nghiem Tuong Diep, Tan Ngoc Pham 等NeurIPS 2023 · 被引用 107 次
