VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis
Linshan Wu, Jiaxin Zhuang, Hao Chen
摘要
Self-Supervised Learning (SSL) has demonstrated promising results in 3D medical image analysis. However, the lack of high-level semantics in pre-training still heavily hinders the performance of downstream tasks. We ob-serve that 3D medical images contain relatively consistent contextual position information, i.e., consistent geometric relations between different organs, which leads to a potential way for us to learn consistent semantic representations in pre-training. In this paper, we propose a simple-yet-effective Volume Contrast (VoCo) framework to leverage the contextual position priors for pre-training. Specif-ically, we first generate a group of base crops from different regions while enforcing feature discrepancy among them, where we employ them as class assignments of dif-ferent regions. Then, we randomly crop sub-volumes and predict them belonging to which class (located at which re-gion) by contrasting their similarity to different base crops, which can be seen as predicting contextual positions of different sub-volumes. Through this pretext task, VoCo implic-itly encodes the contextual position priors into model rep-resentations without the guidance of annotations, enabling us to effectively improve the performance of downstream tasks that require high-level semantics. Extensive exper-imental results on six downstream tasks demonstrate the superior effectiveness of VoCo. Code will be available at httpsu/github.com/luffytls/vo'Co.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert ReasonerWenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng 等AAAI 2026 · 被引用 17 次
- Revisiting 2D Foundation Models for Scalable 3D Medical Image ClassificationHan Liu, Bogdan Georgescu, Yanbo Zhang, Youngjin Yoo 等CVPR 2026 · 被引用 10 次
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT TransformersCris Claessens, Christiaan Viviers, Giacomo D'Amicantonio, Egor Bondarev 等CVPR 2026 · 被引用 6 次
- Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant RepresentationZhaorui Tan, Xi Yang, Tan Pan, Tianyi Liu 等ICCV 2025 · 被引用 5 次
- Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image SegmentationSzymon Plotka, Gizem Mert, Maciej Chrabaszcz, Ewa Szczurek 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Anatomical Invariance Modeling and Semantic Alignment for Self-supervised Learning in 3D Medical Image AnalysisYankai Jiang, Mingze Sun, Heng Guo, Xiaoyu Bai 等ICCV 2023 · 被引用 38 次
- Semantic Information in Contrastive LearningShengjiang Quan, Masahiro Hirano, Yuji YamakawaICCV 2023 · 被引用 4 次
- Autoregressive Sequence Modeling for 3D Medical Image RepresentationSiwen Wang, Churan Wang, Fei Gao, Lixian Su 等AAAI 2025 · 被引用 5 次
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu 等CVPR 2026
- Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labelsJizong Peng, Ping Wang, Christian Desrosiers, Marco PedersoliNeurIPS 2021 · 被引用 80 次
