VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging
Yufan He, Pengfei Guo, Yucheng Tang, Andriy Myronenko, Vishwesh Nath, Ziyue Xu, Dong Yang, Can Zhao, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey
Abstract
Foundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solutions. Specifically, such foundation models should support a full workflow that can actually reduce human effort. Treating 3D medical images as sequences of 2D slices and reusing interactive 2D foundation models seems straightforward, but 2D annotation is too timeconsuming for 3D tasks. Moreover, for large cohort analysis, it's the highly accurate automatic segmentation models that reduce the most human effort. However, these models lack support for interactive corrections and lack zeroshot ability for novel structures, which is a key feature of "foundation". While reusing pre-trained 2D backbones in 3D enhances zero-shot potential, their performance on complex 3D structures still lags behind leading 3D models. To address these issues, we present VISTA3D, Versatile Imaging SegmenTation and Annotation model, that targets to solve all these challenges and requirements with one unified foundation model. VISTA3D is built on top of the well-established 3D segmentation pipeline, and it is the first model to achieve state-of-the-art performance in both 3D automatic (supporting 127 classes) and 3D interactive segmentation, even when compared with top 3D expert models on large and diverse benchmarks. Additionally, VISTA3D's 3D interactive design allows efficient human correction, and a novel 3D supervoxel method that distills 2D pretrained backbones grants VISTA3D top 3D zero-shot performance. We believe the model, recipe, and insights represent a promising step towards a clinically useful 3D foundation model. Code and weights are publicly available at https://github.com/Project-MONAI/VISTA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT TransformersCris Claessens, Christiaan Viviers, Giacomo D'Amicantonio, Egor Bondarev et al.CVPR 2026 · 6 citations
- Foundation VAE for CT Reconstruction, Augmentation, and GenerationQi Chen, Shuhan Ding, Yu Gu, Nan Liu et al.ICML 2026 · 1 citation
- 3DMedAgent: Unified Perception-to-Understanding for 3D Medical AnalysisZiyue Wang, Linghan Cai, Chang Low, Haofeng Liu et al.ICML 2026
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical SegmentationJinpeng Lu, Linghan Cai, Yinda Chen, Guo Tang et al.ICLR 2026
- CP-Agent: Context‑Aware Multimodal Reasoning for Cellular Morphological Profiling under Chemical PerturbationsYuxin Zhang, Yiyao Li, Ping Shu Ho, Simon See et al.ICLR 2026
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionJie Liu, Yixiao Zhang, Jieneng Chen, Junfei Xiao et al.ICCV 2023 · 336 citations
- UniverSeg: Universal Medical Image SegmentationVictor Ion Butoi, Jose Javier Gonzalez Ortiz, Tianyu Ma, Mert R. Sabuncu et al.ICCV 2023 · 163 citations
Related papers
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 155 citations
- Revisiting 2D Foundation Models for Scalable 3D Medical Image ClassificationHan Liu, Bogdan Georgescu, Yanbo Zhang, Youngjin Yoo et al.CVPR 2026 · 10 citations
- vesselFM: A Foundation Model for Universal 3D Blood Vessel SegmentationBastian Wittmann, Yannick Wattenberg, Tamaz Amiranashvili, Suprosanna Shit et al.CVPR 2025
- Interactive Medical Image Segmentation: A Benchmark Dataset and BaselineJunlong Cheng, Bin Fu, Jin Ye, Guoan Wang et al.CVPR 2025
- Generative Medical SegmentationJiayu Huo, Xi Ouyang, Sébastien Ourselin, Rachel SparksAAAI 2025 · 8 citations
