VOILA: Complexity-Aware Universal Segmentation of CT Images by Voxel Interacting with Language
Zishuo Wan, Yu Gao, Wanyuan Pang, Dawei Ding
摘要
Satisfactory progress has been achieved recently in universal segmentation of CT images. Following the success of visionlanguage methods, there is a growing trend towards utilizing text prompts and contrastive learning to develop universal segmentation models. However, there exists a significant imbalance in information density between 3D images and text prompts. Moreover, the standard fully connected layer segmentation approach faces significant challenges in handling multiple classes and exhibits poor generalizability. To address these challenges, we propose the VOxel Interacting with LAnguage method (VOILA) for universal CT image segmentation. Initially, we align voxels and language into a shared representation space and classify voxels on the basis of cosine similarity. Subsequently, we develop the Voxel-Language Interaction framework to mitigate the impact of class imbalance caused by foreground-background discrepancies and variations in target volumes. Furthermore, a Complexity-Aware Sampling method is proposed to focus on region hard to segment, achieved by generating pseudo-heatmaps from a trainable Gaussian mixture distribution. Our results indicate the proposed VOILA is capable to achieve improved performance with reduced parameters and computational cost during training. Furthermore, it demonstrates significant generalizability across diverse datasets without additional finetuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
相关 Paper
- VoxTell: Free-Text Promptable Universal 3D Medical Image SegmentationMaximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee 等CVPR 2026 · 被引用 22 次
- DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image SegmentationQingtao Pan, Wenhao Qiao, Jingjiao Lou, Bing Ji 等AAAI 2025 · 被引用 13 次
- Unsupervised Vision-Language Grammar Induction with Shared Structure ModelingBo Wan, Wenjuan Han, Zilong Zheng, Tinne TuytelaarsICLR 2022 · 被引用 19 次
- Image-Text Co-Decomposition for Text-Supervised Semantic SegmentationJi-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen 等CVPR 2024 · 被引用 6 次
- Universal 3D Shape Matching via Coarse-to-Fine Language GuidanceQinfeng Xiao, Guofeng Mei, Bo Yang, Zhang Liying 等CVPR 2026 · 被引用 1 次
