D^2iT: Dynamic Diffusion Transformer for Accurate Image Generation
Weinan Jia, Mengqi Huang, Nan Chen, Lei Zhang, Zhendong Mao
Abstract
Diffusion models are widely recognized for their ability to generate high-fidelity images. Despite the excellent performance and scalability of the Diffusion Transformer (DiT) architecture, it applies fixed compression across different image regions during the diffusion process, disregarding the naturally varying information densities present in these regions. However, large compression leads to limited local realism, while small compression increases computational complexity and compromises global consistency, ultimately impacting the quality of generated images. To address these limitations, we propose dynamically compressing different image regions by recognizing the importance of different regions, and introduce a novel two-stage framework designed to enhance the effectiveness and efficiency of image generation: (1) Dynamic VAE (DVAE) at first stage employs a hierarchical encoder to encode different image regions at different downsampling rates, tailored to their specific information densities, thereby providing more accurate and natural latent codes for the diffusion process. (2) Dynamic Diffusion Transformer (D 2 iT) at second stage generates images by predicting multi-grained noise, consisting of coarse-grained (less latent code in smooth regions) and fine-grained (more latent codes in detailed regions), through an novel combination of the Dynamic Grain Transformer and the Dynamic Content Transformer. The strategy of combining rough prediction of noise with detailed regions correction achieves a unification of global consistency and local realism. Comprehensive experiments on various generation tasks validate the effectiveness of our approach. Code will be released at https://github . com/jiawn-creator/Dynamic-DiT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f29224f-f1ee-41e2-8356-620de76065a4Cited by top-tier papers5
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and GenerationZhengqiang Zhang, Rongyuan Wu, Lingchen Sun, Lei ZhangNeurIPS 2025 · 8 citations
- Are Image-to-Video Models Good Zero-Shot Image Editors?Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi YangCVPR 2026 · 4 citations
- FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive FocusQiaoqiao Jin, Siming Fu, Dong She, Weinan Jia et al.AAAI 2026 · 1 citation
- LongAnimation: Long Animation Generation with Dynamic Global-Local MemoryNan Chen, Mengqi Huang, Yihao Meng, Zhendong MaoICCV 2025 · 1 citation
- Content-Aware Dynamic Patchification for Efficient Video DiffusionSheng Li, Connelly Barnes, Mamshad Nayeem Rizve, Hongwu Peng et al.CVPR 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion TransformersHaoran You, Connelly Barnes, Yuqian Zhou, Yan Kang et al.CVPR 2025
- DiT-IC: Aligned Diffusion Transformer for Efficient Image CompressionJunqi Shi, Ming Lu, Xingchen Li, Anle Ke et al.CVPR 2026 · 4 citations
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin et al.ICCV 2025
- DDiT: Dynamic Patch Scheduling for Efficient Diffusion TransformersDahye Kim, Deepti Ghadiyaram, Raghudeep GaddeCVPR 2026 · 3 citations
