Revisiting MAE Pre-training for 3D Medical Image Segmentation
Tassilo Wald, Constantin Ulrich, Stanislav Lukyanenko, Andrei Goncharov, Alberto Paderno, Maximilian Miller, Leander Maerkisch, Paul F. Jaeger, Klaus H. Maier-Hein
Abstract
Self-Supervised Learning (SSL) presents an exciting opportunity to unlock the potential of vast, untapped clinical datasets, for various downstream applications that suffer from the scarcity of labeled data. While SSL has revolutionized fields like natural language processing and computer vision, its adoption in 3D medical image computing has been limited by three key pitfalls: Small pre-training dataset sizes, architectures inadequate for 3D medical image analysis, and insufficient evaluation practices. In this paper, we address these issues by i) leveraging a large-scale dataset of 39k 3D brain MRI volumes and ii) using a Residual Encoder U-Net architecture within the state-of-the-art nnU-Net framework. iii) A robust development framework, incorporating 5 development and 8 testing brain MRI segmentation datasets, allowed performance-driven design decisions to optimize the simple concept of Masked Auto Encoders (MAEs) for 3D CNNs. The resulting model not only surpasses previous SSL methods but also outperforms the strong nnU-Net baseline by an average of approximately 3 Dice points setting a new state-of-the-art. Our code and models are made available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c48bbaec-4738-4a9b-afb2-94780a981a3bCited by top-tier papers5
- An OpenMind for 3D Medical Vision Self-supervised LearningTassilo Wald, Constantin Ulrich, Jonathan Suprijadi, Sebastian Ziegler et al.ICCV 2025 · 5 citations
- Disentangling for Transfer: Boosting Limited Modalities via Information-Theoretic Regularization and Cross-Modal ReconstructionZhiyun Zhang, Yan-Jie Zhou, Yujian Hu, Xiyao Ma et al.AAAI 2026
- Modeling the Density of Pixel-level Self-supervised Embeddings for Unsupervised Pathology Segmentation in Medical CTMikhail Goncharov, Eugenia Soboleva, Daniil Ignatyev, Mariia Donskova et al.ICLR 2026
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu et al.CVPR 2026
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical SegmentationJinpeng Lu, Linghan Cai, Yinda Chen, Guo Tang et al.ICLR 2026
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- Contrastive learning of global and local features for medical image segmentation with limited annotationsKrishna Chaitanya, Ertunc Erdil, Neerav Karani, Ender KonukogluNeurIPS 2020 · 714 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
Related papers
- MedGMAE: Gaussian Masked Autoencoders for Medical Volumetric Representation LearningXueming Fu, Fenghe Tang, Rongsheng Wang, Yingtai Li et al.ICLR 2026
- Modeling the Probabilistic Distribution of Unlabeled Data for One-shot Medical Image SegmentationYuhang Ding, Xin Yu, Yi YangAAAI 2021 · 42 citations
- Autoregressive Sequence Modeling for 3D Medical Image RepresentationSiwen Wang, Churan Wang, Fei Gao, Lixian Su et al.AAAI 2025 · 5 citations
- Cabld: Contrast-Agnostic Brain Landmark Detection With Consistency-Based RegularizationSoorena Salari, Arash Harirpoush, Hassan Rivaz, Yiming XiaoICCV 2025 · 1 citation
- Masked LoGoNet: Fast and Accurate 3D Image Analysis for Medical DomainAmin Karimi Monsefi, Payam Karisani, Mengxi Zhou, Stacey Choi et al.KDD 2024 · 2 citations
