Foundation VAE for CT Reconstruction, Augmentation, and Generation
Qi Chen, Shuhan Ding, Yu Gu, Nan Liu, Jiang Bian, Alan Yuille, Zongwei Zhou, Jingjing Fu
Abstract
Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserving clinically relevant structure. However, training CT-specific VAEs from scratch or heavily fine-tuning them incurs substantial computational and engineering cost, and often degrades under heterogeneous scanners, protocols, and diseases. This paper makes a progressive stride toward training-free medical VAEs by leveraging a critical observation: a single Foundation VAE, pretrained at scale on natural images and videos, can serve as a unified interface for CT Reconstruction, Augmentation, and Generation. With both encoder and decoder frozen, the Foundation VAE reconstructs CT volumes with preserved anatomy while suppressing acquisition noise; training segmentation models on these reconstructions improves surface accuracy by 3.9% NSD on average for pancreatic tumor and lung tumor. Within the same Foundation VAE latent space, a conditional latent diffusion model achieves 3.9% lower average FVD with 36.2% higher CT CLIP score, and improves multi-disease generation faithfulness across 18 types by 2.76% AUC. These results demonstrate Foundation VAEs as a practical interface for scalable CT representation reuse and faithful CT generation. Our code and demo are available at https://github.com/qic999/Foundation-VAE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsSijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang et al.NeurIPS 2024 · 82 citations
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical ImagingIbrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Hadrien Reynaud et al.NeurIPS 2025 · 23 citations
- LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion ModelsYu Cheng, Fajie YuanICCV 2025 · 1 citation
- Towards Generalizable Tumor SynthesisQi Chen, Xiaoxi Chen, Haorui Song, Zhiwei Xiong et al.CVPR 2024
- VISTA3D: A Unified Segmentation Foundation Model For 3D Medical ImagingYufan He, Pengfei Guo, Yucheng Tang, Andriy Myronenko et al.CVPR 2025
Related papers
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu et al.CVPR 2026
- GuideGen: A Text-Guided Framework for Paired Full-torso Anatomy and CT Volume GenerationLinrui Dai, Rongzhao Zhang, Yongrui Yu, Xiaofan ZhangAAAI 2026
- Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion ModelsYankai Jiang, Peng Zhang, Donglin Yang, Yuan Tian et al.CVPR 2025
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 155 citations
- TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion ModelZhenkai Zhang, Krista A. Ehinger, Tom DrummondAAAI 2025
