Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
Jiayun Jin, Haolong Chai, Xueying Huang, Xiaoqing Guo, Zengwei Zheng, Zhan Zhou, Junmei Wang, Xinyu Wang, Jie Liu, Binbin Zhou
摘要
Ultrasound imaging is widely used in clinical diagnostics due to its real-time capability and radiation-free nature. However, existing vision-language pre-training models, such as CLIP, are primarily designed for modalities like CT and MRI, and are difficult to directly apply to ultrasound data, which exhibit heterogeneous anatomical structures and diverse diagnostic attributes.To bridge this gap, we construct US-365K, a large-scale ultrasound image–text dataset containing 365k paired samples across 52 anatomical categories. We establish Ultrasonographic Diagnostic Taxonomy (UDT) containing two hierarchical knowledge frameworks, Ultrasonographic Hierarchical Anatomical Taxonomy (UHAT) and Ultrasonographic Diagnostic Attribute Framework (UDAF). UHAT standardizes anatomical organization, and UDAF formalizes nine diagnostic dimensions, including body system, organ, diagnosis, shape, margins, echogenicity, internal characteristics, posterior acoustic phenomena, and vascularity.Building upon these foundations, we propose Ultrasound-CLIP, a semantic-aware contrastive learning framework that introduces semantic soft labels and semantic loss to refine sample discrimination. Moreover, we construct a heterogeneous graph modality derived from UDAF's textual representations, enabling structured reasoning over lesion–attribute relations.Extensive experiments with patient-level data splitting demonstrate that our approach achieves state-of-the-art performance on classification and retrieval benchmarks built upon US-365K, while also delivering strong generalization to zero-shot, linear probing, and fine-tuning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionJie Liu, Yixiao Zhang, Jieneng Chen, Junfei Xiao 等ICCV 2023 · 被引用 336 次
相关 Paper
- Boosting Medical Visual Understanding From Multi-Granular Language LearningZihan Li, Yiqing Wang, Sina Farsiu, Paul KinahanICLR 2026 · 被引用 6 次
- Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologySiyuan Yan, Ming Hu, Yiwen Jiang, Xieji Li 等ICCV 2025 · 被引用 7 次
- Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsYinuo Jing, Chunyu Wang, Ruxu Zhang, Kongming Liang 等ACM MM 2023 · 被引用 6 次
- SEMC: Structure-Enhanced Mixture-of-Experts Contrastive Learning for Ultrasound Standard Plane RecognitionQing Cai, Guihao Yan, Fan Zhang, Cheng Zhang 等AAAI 2026
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image UnderstandingZhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang 等ICLR 2025
