Representation Alignment for Diffusion Transformers without External Components
Dengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang, Haoyu Wang, Wei Wei, Guang Dai, Yanning Zhang, Jingdong Wang
摘要
Recent studies have demonstrated that learning a meaningful internal representation can accelerate generative training. However, existing approaches necessitate to either introduce an off-the-shelf external representation task or rely on a large-scale, pre-trained external representation encoder to provide representation guidance during the training process. In this study, we posit that the unique discriminative process inherent to diffusion transformers enables them to offer such guidance without requiring external representation components. We propose Self-Representation Alignment (SRA), a simple yet effective method that obtains representation guidance using the internal representations of learned diffusion transformer. SRA aligns the latent representation of the diffusion transformer in the earlier layer conditioned on higher noise to that in the later layer conditioned on lower noise to progressively enhance the overall representation learning during only the training process. Experimental results indicate that applying SRA to DiTs and SiTs yields consistent performance improvements, and largely outperforms approaches relying on auxiliary representation task. Our approach achieves performance comparable to methods that are dependent on an external pre-trained representation encoder, which demonstrates the feasibility of acceleration with ‡ This work was completed during the internship at SGIT AI Lab, State Grid Corporation of China. * Corresponding author. † Project lead.
In the following part, we use term 'student' to represent the trainable model and 'teacher' to represent the EMA model in SRA for simplicity.
- For SiTs, tmax = 1 and for DiTs, tmax = 1000. In practice, we truncate (t -k) to 0 if it is less than 0.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You ThinkGe Wu, Shen Zhang, Ruijing Shi, Shanghua Gao 等NeurIPS 2025 · 被引用 102 次
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 被引用 102 次
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng 等NeurIPS 2025 · 被引用 98 次
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang 等CVPR 2026 · 被引用 59 次
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 等CVPR 2026 · 被引用 36 次
它引用的顶会 Paper54
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
相关 Paper
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You ThinkSihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong 等ICLR 2025
- SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion TrainingMengmeng Wang, Dengyang Jiang, Liuzhuozheng Li, Yucheng Lin 等CVPR 2026 · 被引用 9 次
- DiverseDiT: Towards Diverse Representation Learning in Diffusion TransformersMengping Yang, Zhiyu Tan, Binglei Li, Xiaomeng Yang 等CVPR 2026 · 被引用 4 次
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingZiqiao Wang, Wangbo Zhao, Yuhao Zhou, Zekai Li 等NeurIPS 2025 · 被引用 37 次
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 被引用 288 次
