Towards Text-Mask Consistency in Medical Image Segmentation
Jie Gui, HangTu, Wen Sha, Xiuquan Du
摘要
Vision-language models for medical image segmentation often produce masks that conflict with the accompanying text, especially under multi-site/multi-lesion descriptions. We trace this failure to two factors: (i) highly templated and repetitive clinical language causes one-to-one hard contrastive learning to yield numerous false negatives, weakening cross-modal alignment; and (ii) predominantly vision-driven, one-way cross-attention lacks a language-dominant, spatially aware pathway, hindering effective injection of textual semantics into the spatial visual domain. To this end, we propose Consistency-enhanced Two-stage Segmentation (C2Seg). In the pretraining stage, Cluster-aware Contrastive Learning uses a frozen strong baseline to construct an intra-batch text similarity matrix as soft labels, thereby alleviating false negative conflicts and producing more discriminative visual representations. In the fusion stage, we introduce a Bidirectional Complementary Attention Module, where each modality dominates attention along its own path, fostering deep interaction and structural consistency between visual and textual representations. In order to enhance the expressive power of multimodal features, we further adopt KAN-based Attention Gating. Without updating the language encoder, our approach significantly improves text-mask consistency and segmentation accuracy on four public medical imaging datasets. * Corresponding author. RELATED WORK Vision Language Models. In recent years, the success of general vision-language pretraining models has driven multimodal research in the medical domain. For example, Zhang et al. (2025b)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual RepresentationsDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet 等ICCV 2021 · 被引用 542 次
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image RecognitionShih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena YeungICCV 2021 · 被引用 516 次
相关 Paper
- DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image SegmentationQingtao Pan, Wenhao Qiao, Jingjiao Lou, Bing Ji 等AAAI 2025 · 被引用 13 次
- PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsPujin Cheng, Li Lin, Junyan Lyu, Yijin Huang 等ICCV 2023 · 被引用 91 次
- Boosting Medical Visual Understanding From Multi-Granular Language LearningZihan Li, Yiqing Wang, Sina Farsiu, Paul KinahanICLR 2026 · 被引用 6 次
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionPeng Zhang, Zhihui Lai, Wenting Chen, Xu Wu 等AAAI 2026
- CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentSajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo 等CVPR 2024
