Anatomical Invariance Modeling and Semantic Alignment for Self-supervised Learning in 3D Medical Image Analysis
Yankai Jiang, Mingze Sun, Heng Guo, Xiaoyu Bai, Ke Yan, Le Lu, Minfeng Xu
Abstract
Self-supervised learning (SSL) has recently achieved promising performance for 3D medical image analysis tasks. Most current methods follow existing SSL paradigm originally designed for photographic or natural images, which cannot explicitly and thoroughly exploit the intrinsic similar anatomical structures across varying medical images. This may in fact degrade the quality of learned deep representations by maximizing the similarity among features containing spatial misalignment information and different anatomical semantics. In this work, we propose a new self-supervised learning framework, namely Alice, that explicitly fulfills Anatomical invariance modeling and semantic alignment via elaborately combining discriminative and generative objectives. Alice introduces a new contrastive learning strategy which encourages the similarity between views that are diversely mined but with consistent high-level semantics, in order to learn invariant anatomical features. Moreover, we design a conditional anatomical feature alignment module to complement corrupted embeddings with globally matched semantics and inter-patch topology information, conditioned by the distribution of local image content, which permits to create better contrastive pairs. Our extensive quantitative experiments on three 3D medical image analysis tasks demonstrate and validate the performance superiority of Alice, surpassing the previous best SSL counterpart methods and showing promising ability for united representation learning. Codes are available at https://github.com/alibaba- damo-academy/alice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisLinshan Wu, Jiaxin Zhuang, Hao ChenCVPR 2024 · 60 citations
- Continual Self-Supervised Learning: Towards Universal Multi-Modal Medical Data Representation LearningYiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen et al.CVPR 2024 · 30 citations
- Representing Part-Whole Hierarchies in Foundation Models by Learning Localizability, Composability, and Decomposability from Anatomy via Self-SupervisionMohammad Reza Hosseinzadeh Taher, Michael B. Gotway, Jianming LiangCVPR 2024 · 12 citations
- Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant RepresentationZhaorui Tan, Xi Yang, Tan Pan, Tianyi Liu et al.ICCV 2025 · 5 citations
- Structure-Aware Semantic Discrepancy and Consistency for 3D Medical Image Self-Supervised LearningTan Pan, Zhaorui Tan, Kaiyu Guo, Dongli Xu et al.ICCV 2025 · 2 citations
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Autoregressive Sequence Modeling for 3D Medical Image RepresentationSiwen Wang, Churan Wang, Fei Gao, Lixian Su et al.AAAI 2025 · 5 citations
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report GenerationLongzhen Yang, Zhangkai Ni, Ying Wen, Yihang Liu et al.ACM MM 2025
- Pseudo-Label Guided Contrastive Learning for Semi-Supervised Medical Image SegmentationHritam Basak, Zhaozheng YinCVPR 2023
- Contrastive learning of global and local features for medical image segmentation with limited annotationsKrishna Chaitanya, Ertunc Erdil, Neerav Karani, Ender KonukogluNeurIPS 2020 · 714 citations
- Geometric Visual Similarity Learning in 3D Medical Image Self-Supervised Pre-trainingYuting He, Guanyu Yang, Rongjun Ge, Yang Chen et al.CVPR 2023
