DDT: Decoupled Diffusion Transformer
Shuai Wang, Zhi Tian, Weilin Huang, Limin Wang
摘要
Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs to extract the lower-frequency semantic component and then decode the higher frequency with identical modules. This scheme creates an inherent optimization dilemma: encoding low-frequency semantics necessitates reducing high-frequency components, creating tension between semantic encoding and high-frequency decoding. To resolve this challenge, we propose a new ddtDecoupled ddtDiffusion ddtTransformer (ddtDDT), with a decoupled design of a dedicated condition encoder for semantic extraction alongside a specialized velocity decoder. Our experiments reveal that a more substantial encoder yields performance improvements as model size increases. For ImageNet , Our DDT-XL/2 achieves a new state-of-the-art performance of 1.31 FID (nearly faster training convergence compared to previous diffusion transformers). For ImageNet , Our DDT-XL/2 achieves a new state-of-the-art FID of 1.28. Additionally, as a beneficial by-product, our decoupled architecture enhances inference speed by enabling the sharing self-condition between adjacent denoising steps. To minimize performance degradation, we propose a novel statistical dynamic programming approach to identify optimal sharing strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Representation Alignment for Diffusion Transformers without External ComponentsDengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang 等ICLR 2026 · 被引用 532 次
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 被引用 288 次
- Improved Mean Flows: On the Challenges of Fastforward Generative ModelsZhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman 等CVPR 2026 · 被引用 116 次
- What matters for Representation Alignment: Global Information or Spatial Structure?Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng 等ICLR 2026 · 被引用 84 次
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng 等CVPR 2026 · 被引用 82 次
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang 等CVPR 2026 · 被引用 59 次
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated SamplingKyungmin Lee, Sihyun Yu, Jinwoo ShinICLR 2026 · 被引用 18 次
- Language-Guided Image Tokenization for GenerationKaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross 等CVPR 2025
- Denoising Token Prediction in Masked Autoregressive ModelsTing Yao, Yehao Li, Yingwei Pan, Zhaofan Qiu 等ICCV 2025 · 被引用 2 次
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderXiaojie Li, Yang Zhao, Ming Li, Yancheng Zhang 等CVPR 2026
