VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis
Linshan Wu, Jiaxin Zhuang, Hao Chen
Abstract
Self-Supervised Learning (SSL) has demonstrated promising results in 3D medical image analysis. However, the lack of high-level semantics in pre-training still heavily hinders the performance of downstream tasks. We ob-serve that 3D medical images contain relatively consistent contextual position information, i.e., consistent geometric relations between different organs, which leads to a potential way for us to learn consistent semantic representations in pre-training. In this paper, we propose a simple-yet-effective Volume Contrast (VoCo) framework to leverage the contextual position priors for pre-training. Specif-ically, we first generate a group of base crops from different regions while enforcing feature discrepancy among them, where we employ them as class assignments of dif-ferent regions. Then, we randomly crop sub-volumes and predict them belonging to which class (located at which re-gion) by contrasting their similarity to different base crops, which can be seen as predicting contextual positions of different sub-volumes. Through this pretext task, VoCo implic-itly encodes the contextual position priors into model rep-resentations without the guidance of annotations, enabling us to effectively improve the performance of downstream tasks that require high-level semantics. Extensive exper-imental results on six downstream tasks demonstrate the superior effectiveness of VoCo. Code will be available at httpsu/github.com/luffytls/vo'Co.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8fb123c-2c63-4ef6-a7a0-231886a33377Cited by top-tier papers26
- Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert ReasonerWenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng et al.AAAI 2026 · 17 citations
- Revisiting 2D Foundation Models for Scalable 3D Medical Image ClassificationHan Liu, Bogdan Georgescu, Yanbo Zhang, Youngjin Yoo et al.CVPR 2026 · 10 citations
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT TransformersCris Claessens, Christiaan Viviers, Giacomo D'Amicantonio, Egor Bondarev et al.CVPR 2026 · 6 citations
- Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant RepresentationZhaorui Tan, Xi Yang, Tan Pan, Tianyi Liu et al.ICCV 2025 · 5 citations
- Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image SegmentationSzymon Plotka, Gizem Mert, Maciej Chrabaszcz, Ewa Szczurek et al.NeurIPS 2025 · 5 citations
Builds on24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Anatomical Invariance Modeling and Semantic Alignment for Self-supervised Learning in 3D Medical Image AnalysisYankai Jiang, Mingze Sun, Heng Guo, Xiaoyu Bai et al.ICCV 2023 · 38 citations
- Semantic Information in Contrastive LearningShengjiang Quan, Masahiro Hirano, Yuji YamakawaICCV 2023 · 4 citations
- Autoregressive Sequence Modeling for 3D Medical Image RepresentationSiwen Wang, Churan Wang, Fei Gao, Lixian Su et al.AAAI 2025 · 5 citations
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu et al.CVPR 2026
- Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labelsJizong Peng, Ping Wang, Christian Desrosiers, Marco PedersoliNeurIPS 2021 · 80 citations
