Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
Jiayun Jin, Haolong Chai, Xueying Huang, Xiaoqing Guo, Zengwei Zheng, Zhan Zhou, Junmei Wang, Xinyu Wang, Jie Liu, Binbin Zhou
Abstract
Ultrasound imaging is widely used in clinical diagnostics due to its real-time capability and radiation-free nature. However, existing vision-language pre-training models, such as CLIP, are primarily designed for modalities like CT and MRI, and are difficult to directly apply to ultrasound data, which exhibit heterogeneous anatomical structures and diverse diagnostic attributes.To bridge this gap, we construct US-365K, a large-scale ultrasound image–text dataset containing 365k paired samples across 52 anatomical categories. We establish Ultrasonographic Diagnostic Taxonomy (UDT) containing two hierarchical knowledge frameworks, Ultrasonographic Hierarchical Anatomical Taxonomy (UHAT) and Ultrasonographic Diagnostic Attribute Framework (UDAF). UHAT standardizes anatomical organization, and UDAF formalizes nine diagnostic dimensions, including body system, organ, diagnosis, shape, margins, echogenicity, internal characteristics, posterior acoustic phenomena, and vascularity.Building upon these foundations, we propose Ultrasound-CLIP, a semantic-aware contrastive learning framework that introduces semantic soft labels and semantic loss to refine sample discrimination. Moreover, we construct a heterogeneous graph modality derived from UDAF's textual representations, enabling structured reasoning over lesion–attribute relations.Extensive experiments with patient-level data splitting demonstrate that our approach achieves state-of-the-art performance on classification and retrieval benchmarks built upon US-365K, while also delivering strong generalization to zero-shot, linear probing, and fine-tuning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68ce6b20-497f-4d38-ab25-ea51e63f87c0Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionJie Liu, Yixiao Zhang, Jieneng Chen, Junfei Xiao et al.ICCV 2023 · 336 citations
Related papers
- Boosting Medical Visual Understanding From Multi-Granular Language LearningZihan Li, Yiqing Wang, Sina Farsiu, Paul KinahanICLR 2026 · 6 citations
- Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologySiyuan Yan, Ming Hu, Yiwen Jiang, Xieji Li et al.ICCV 2025 · 7 citations
- Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsYinuo Jing, Chunyu Wang, Ruxu Zhang, Kongming Liang et al.ACM MM 2023 · 6 citations
- SEMC: Structure-Enhanced Mixture-of-Experts Contrastive Learning for Ultrasound Standard Plane RecognitionQing Cai, Guihao Yan, Fan Zhang, Cheng Zhang et al.AAAI 2026
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image UnderstandingZhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang et al.ICLR 2025
