Representation Alignment for Diffusion Transformers without External Components
Dengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang, Haoyu Wang, Wei Wei, Guang Dai, Yanning Zhang, Jingdong Wang
Abstract
Recent studies have demonstrated that learning a meaningful internal representation can accelerate generative training. However, existing approaches necessitate to either introduce an off-the-shelf external representation task or rely on a large-scale, pre-trained external representation encoder to provide representation guidance during the training process. In this study, we posit that the unique discriminative process inherent to diffusion transformers enables them to offer such guidance without requiring external representation components. We propose Self-Representation Alignment (SRA), a simple yet effective method that obtains representation guidance using the internal representations of learned diffusion transformer. SRA aligns the latent representation of the diffusion transformer in the earlier layer conditioned on higher noise to that in the later layer conditioned on lower noise to progressively enhance the overall representation learning during only the training process. Experimental results indicate that applying SRA to DiTs and SiTs yields consistent performance improvements, and largely outperforms approaches relying on auxiliary representation task. Our approach achieves performance comparable to methods that are dependent on an external pre-trained representation encoder, which demonstrates the feasibility of acceleration with ‡ This work was completed during the internship at SGIT AI Lab, State Grid Corporation of China. * Corresponding author. † Project lead.
In the following part, we use term 'student' to represent the trainable model and 'teacher' to represent the EMA model in SRA for simplicity.
- For SiTs, tmax = 1 and for DiTs, tmax = 1000. In practice, we truncate (t -k) to 0 if it is less than 0.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fdee441-068e-4794-96eb-b01fe4f13cb0Cited by top-tier papers64
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You ThinkGe Wu, Shen Zhang, Ruijing Shi, Shanghua Gao et al.NeurIPS 2025 · 102 citations
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 102 citations
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng et al.NeurIPS 2025 · 98 citations
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang et al.CVPR 2026 · 59 citations
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou et al.CVPR 2026 · 36 citations
Builds on54
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You ThinkSihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong et al.ICLR 2025
- SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion TrainingMengmeng Wang, Dengyang Jiang, Liuzhuozheng Li, Yucheng Lin et al.CVPR 2026 · 9 citations
- DiverseDiT: Towards Diverse Representation Learning in Diffusion TransformersMengping Yang, Zhiyu Tan, Binglei Li, Xiaomeng Yang et al.CVPR 2026 · 4 citations
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingZiqiao Wang, Wangbo Zhao, Yuhao Zhou, Zekai Li et al.NeurIPS 2025 · 37 citations
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 288 citations
