Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-Training
Weiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang, Zeli Chen, Xi Li, Le Lu, Xianghua Ye, Tingbo Liang, Qi Zhang, Ling Zhang
Abstract
Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a high SNR presents a semantic density gap, leading to visual alignment bias. In this paper, we propose boosting vision semantic density to improve alignment effectiveness. On one hand, we enhance visual semantics through disease-level vision contrastive learning, which strengthens the model's ability to differentiate between normal and abnormal samples for each anatomical structure. On the other hand, we introduce an anatomical normality modeling method to model the distribution of normal samples for each anatomy, leveraging VQ-VAE for reconstructing normal vision embeddings in the latent space. This process amplifies abnormal signals by leveraging distribution shifts in abnormal samples, enhancing the model's perception and discrimination of abnormal attributes. The enhanced visual representation effectively captures the diagnostic-relevant semantics, facilitating more efficient and accurate alignment with the diagnostic report. We conduct extensive experiments on two chest CT datasets, CT-RATE and Rad-ChestCT, and an abdominal CT dataset, MedVL-CT69K, and comprehensively evaluate the diagnosis performance across multiple tasks in the chest and abdominal CT scenarios, achieving state-of-theart zero-shot performance. Notably, our method achieved an average AUC of 84.9% across 54 diseases in 15 organs, significantly surpassing existing methods. Additionally, we demonstrate the superior transfer learning capabilities of our pre-trained model. Code is available at https:// github.com/alibaba-damo-academy/ViSD-Boost
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae393c30-8cc0-401a-a15e-1df9347072c8Cited by top-tier papers2
- Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer AnalysisYuanzhe Li, Hao Chen, Rui Yin, Juyan Ba et al.CVPR 2026 · 3 citations
- Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed TomographyBowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang et al.ICML 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang et al.CVPR 2022 · 839 citations
Related papers
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image UnderstandingZhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang et al.ICLR 2025
- Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training FrameworkVu Minh Hieu Phan, Yutong Xie, Yuankai Qi, Lingqiao Liu et al.CVPR 2024 · 17 citations
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionPeng Zhang, Zhihui Lai, Wenting Chen, Xu Wu et al.AAAI 2026
- MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisChaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang et al.ICCV 2023 · 205 citations
- CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentSajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo et al.CVPR 2024
