U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
Yuchuan Tian, Zhijun Tu, Hanting Chen, Jie Hu, Chao Xu, Yunhe Wang
Abstract
Diffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of transformer blocks, DiTs demonstrate competitive performance and good scalability; but meanwhile, the abandonment of U-Net by DiTs and their following improvements is worth rethinking. To this end, we conduct a simple toy experiment by comparing a U-Net architectured DiT with an isotropic one. It turns out that the U-Net architecture only gain a slight advantage amid the U-Net inductive bias, indicating potential redundancies within the U-Net-style DiT. Inspired by the discovery that U-Net backbone features are low-frequency-dominated, we perform token downsampling on the query-key-value tuple for self-attention that bring further improvements despite a considerable amount of reduction in computation. Based on self-attention with downsampled tokens, we propose a series of U-shaped DiTs (U-DiTs) in the paper and conduct extensive experiments to demonstrate the extraordinary performance of U-DiT models. The proposed U-DiT could outperform DiT-XL/2 with only 1/6 of its computation cost. Codes are available at https://github.com/YuchuanTian/U-DiT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf219428-860e-464b-9deb-83fc9b299d66Cited by top-tier papers25
- Align Your Flow: Scaling Continuous-Time Flow Map DistillationAmirmojtaba Sabour, Sanja Fidler, Karsten KreisNeurIPS 2025 · 91 citations
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang et al.CVPR 2026 · 59 citations
- U-REPA: Aligning Diffusion U-Nets to ViTsYuchuan Tian, Hanting Chen, Mengyu Zheng, Yuchen Liang et al.NeurIPS 2025 · 30 citations
- FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion GenerationZeyu Zhang, Yiran Wang, Danning Li, Dong Gong et al.NeurIPS 2025 · 12 citations
- Learning to Integrate Diffusion ODEs by Averaging the DerivativesWenze Liu, Xiangyu YueNeurIPS 2025 · 9 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Effective Diffusion Transformer Architecture for Image Super-ResolutionKun Cheng, Lei Yu, Zhijun Tu, Xiao He et al.AAAI 2025 · 26 citations
- DiT-IC: Aligned Diffusion Transformer for Efficient Image CompressionJunqi Shi, Ming Lu, Xingchen Li, Anle Ke et al.CVPR 2026 · 4 citations
- Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion ModelQuan Dao, Dimitris N. MetaxasCVPR 2026
- All are Worth Words: A ViT Backbone for Diffusion ModelsFan Bao, Shen Nie, Kaiwen Xue, Yue Cao et al.CVPR 2023
