Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation
Zhenkai Zhang, Markus Hiller, Krista A. Ehinger, Tom Drummond
摘要
Generating high-resolution 3D CT volumes with fine details remains challenging due to substantial computational demands and optimization difficulties inherent to existing generative models. In this paper, we propose the Pixel-Level Residual Diffusion Transformer (PRDiT), a scalable generative framework that synthesizes high-quality 3D medical volumes directly at voxel-level. PRDiT introduces a two-stage training architecture comprising 1) a local denoiser in the form of an MLP-based blind estimator operating on overlapping 3D patches to separate lowfrequency structures efficiently, and 2) a global residual diffusion transformer employing memory-efficient attention to model and refine high-frequency residuals across entire volumes. This coarse-to-fine modeling strategy simplifies optimization, enhances training stability, and effectively preserves subtle structures without the limitations of an autoencoder bottleneck. Extensive experiments conducted on the LIDC-IDRI and RAD-ChestCT datasets demonstrate that PRDiT consistently outperforms state-of-the-art models, such as HA-GAN, 3D LDM and WDM-3D, achieving significantly lower 3D FID, MMD and Wasserstein distance scores 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape GenerationShentong Mo, Enze Xie, Ruihang Chu, Lanqing Hong 等NeurIPS 2023 · 被引用 157 次
- TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion ModelZhenkai Zhang, Krista A. Ehinger, Tom DrummondAAAI 2025
- Analyzing and Improving the Image Quality of StyleGANTero Karras, Samuli Laine, Miika Aittala, Janne Hellsten 等CVPR 2020
相关 Paper
- Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion TransformersKatherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham 等ICML 2024 · 被引用 98 次
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao 等CVPR 2026 · 被引用 1 次
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng 等CVPR 2026 · 被引用 82 次
- Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume GenerationDelin An, Chaoli WangCVPR 2026
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang 等CVPR 2026 · 被引用 46 次
