MedGMAE: Gaussian Masked Autoencoders for Medical Volumetric Representation Learning
Xueming Fu, Fenghe Tang, Rongsheng Wang, Yingtai Li, Lixia Han, Jian Lu, Zihang Jiang, S Kevin Zhou
Abstract
Self-supervised pre-training has emerged as a critical paradigm for learning transferable representations from unlabeled medical volumetric data. Masked autoencoder based methods have garnered significant attention, yet their application to volumetric medical image faces fundamental limitations from the discrete voxellevel reconstruction objective, which neglects comprehensive anatomical structure continuity. To address this challenge, We propose MedGMAE, a novel framework that replaces traditional voxel reconstruction with 3D Gaussian primitives reconstruction as new perspectives on representation learning. Our approach learns to predict complete sets of 3D Gaussian parameters as semantic abstractions to represent the entire 3D volume, from sparse visible image patches. MedGMAE demonstrates dual utility across medical imaging applications. For representation learning, sparse Gaussian prediction produces superior encoder representations that outperform traditional MAE baselines on downstream segmentation, classification, and registration tasks. For volumetric reconstruction, the Gaussian decoder leverages pretrained anatomical priors to accelerate 3D CT volume reconstruction convergence. Extensive experiments across multiple medical imaging datasets demonstrate that our approach achieves superior performance, establishing a new framework for medical image pre-training. The code will be available in https://github.com/windrise/MedGMAE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91772ff0-1c3e-40ba-ae96-73073d39bd5fBuilds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- R2-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic ReconstructionRuyi Zha, Tao Jun Lin, Yuanhao Cai, Jiwen Cao et al.NeurIPS 2024 · 99 citations
Related papers
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu et al.CVPR 2026
- Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative ModelsBenjamin Eckart, Wentao Yuan, Chao Liu, Jan KautzCVPR 2021
- Revisiting MAE Pre-training for 3D Medical Image SegmentationTassilo Wald, Constantin Ulrich, Stanislav Lukyanenko, Andrei Goncharov et al.CVPR 2025
- Tracking by Predicting 3-D Gaussians Over TimeTanish Baranwal, Himanshu Gaurav Singh, Jathushan Rajasegaran, Jitendra MalikCVPR 2026 · 1 citation
- VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisLinshan Wu, Jiaxin Zhuang, Hao ChenCVPR 2024 · 60 citations
