TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model
Zhenkai Zhang, Krista A. Ehinger, Tom Drummond
Abstract
We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplaneaware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128 × 128 × 128, Pancreas 256 × 256 × 256, and Colon 512×512×512, demonstrate outstanding results. We utilized MSE and SSIM to assess reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons with existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar sized latent spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- Autodecoding Latent 3D Diffusion ModelsEvangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang et al.NeurIPS 2023 · 65 citations
- Class-Balancing Diffusion ModelsYiming Qin, Huangjie Zheng, Jiangchao Yao, Mingyuan Zhou et al.CVPR 2023
- 3D Neural Field Generation Using Triplane DiffusionJ. Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner et al.CVPR 2023
Related papers
- GuideGen: A Text-Guided Framework for Paired Full-torso Anatomy and CT Volume GenerationLinrui Dai, Rongzhao Zhang, Yongrui Yu, Xiaofan ZhangAAAI 2026
- Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion ModelsSuhyeon Lee, Hyungjin Chung, Minyoung Park, Jonghyuk Park et al.ICCV 2023 · 72 citations
- DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelYinghao Xu, Hao Tan, Fujun Luan, Sai Bi et al.ICLR 2024 · 234 citations
- DiffusionBlend: Learning 3D Image Prior through Position-aware Diffusion Score Blending for 3D Computed Tomography ReconstructionBowen Song, Jason Hu, Zhaoxu Luo, Jeffrey A. Fessler et al.NeurIPS 2024 · 20 citations
- Foundation VAE for CT Reconstruction, Augmentation, and GenerationQi Chen, Shuhan Ding, Yu Gu, Nan Liu et al.ICML 2026 · 1 citation
