Foundation VAE for CT Reconstruction, Augmentation, and Generation
Qi Chen, Shuhan Ding, Yu Gu, Nan Liu, Jiang Bian, Alan Yuille, Zongwei Zhou, Jingjing Fu
摘要
Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserving clinically relevant structure. However, training CT-specific VAEs from scratch or heavily fine-tuning them incurs substantial computational and engineering cost, and often degrades under heterogeneous scanners, protocols, and diseases. This paper makes a progressive stride toward training-free medical VAEs by leveraging a critical observation: a single Foundation VAE, pretrained at scale on natural images and videos, can serve as a unified interface for CT Reconstruction, Augmentation, and Generation. With both encoder and decoder frozen, the Foundation VAE reconstructs CT volumes with preserved anatomy while suppressing acquisition noise; training segmentation models on these reconstructions improves surface accuracy by 3.9% NSD on average for pancreatic tumor and lung tumor. Within the same Foundation VAE latent space, a conditional latent diffusion model achieves 3.9% lower average FVD with 36.2% higher CT CLIP score, and improves multi-disease generation faithfulness across 18 types by 2.76% AUC. These results demonstrate Foundation VAEs as a practical interface for scalable CT representation reuse and faithful CT generation. Our code and demo are available at https://github.com/qic999/Foundation-VAE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsSijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang 等NeurIPS 2024 · 被引用 82 次
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical ImagingIbrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Hadrien Reynaud 等NeurIPS 2025 · 被引用 23 次
- LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion ModelsYu Cheng, Fajie YuanICCV 2025 · 被引用 1 次
- Towards Generalizable Tumor SynthesisQi Chen, Xiaoxi Chen, Haorui Song, Zhiwei Xiong 等CVPR 2024
- VISTA3D: A Unified Segmentation Foundation Model For 3D Medical ImagingYufan He, Pengfei Guo, Yucheng Tang, Andriy Myronenko 等CVPR 2025
相关 Paper
- Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-SupervisionYunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu 等CVPR 2026
- GuideGen: A Text-Guided Framework for Paired Full-torso Anatomy and CT Volume GenerationLinrui Dai, Rongzhao Zhang, Yongrui Yu, Xiaofan ZhangAAAI 2026
- Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion ModelsYankai Jiang, Peng Zhang, Donglin Yang, Yuan Tian 等CVPR 2025
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 被引用 155 次
- TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion ModelZhenkai Zhang, Krista A. Ehinger, Tom DrummondAAAI 2025
