What matters for Representation Alignment: Global Information or Spatial Structure?
Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng, Richard Zhang, Eli Shechtman, Saining Xie
摘要
Representation alignment (REPA) guides generative training by distilling representations from a strong, pretrained vision encoder to intermediate diffusion features. We investigate a fundamental question: what aspect of the target representation matters for generation, its global semantic information (e.g., measured by ImageNet-1K accuracy) or its spatial structure (i.e. pairwise cosine similarity between patch tokens)? Prevalent wisdom holds that stronger global semantic performance leads to better generation as a target representation. To study this, we first perform a large-scale empirical analysis across 27 different vision encoders and different model scales. The results are surprising; spatial structure, rather than global performance, drives the generation performance of a target representation. To further study this, we introduce two straightforward modifications, which specifically accentuate the transfer of spatial information. We replace the standard MLP projection layer in REPA with a simple convolution layer and introduce a spatial normalization layer for the external representation. Surprisingly, our simple method (implemented in 4 lines of code), termed iREPA, consistently improves convergence speed of REPA, across a diverse set of vision encoders, model sizes, and training variants (such as REPA, REPA-E, Meanflow, JiT etc). %, etc. Our work motivates revisiting the fundamental working mechanism of representational alignment and how it can be leveraged for improved training of generative models. The code and project page are available at https://end2end-diffusion.github.io/irepa
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi 等ICLR 2026 · 被引用 14 次
- Self-Supervised Flow Matching for Scalable Multi-Modal SynthesisHila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell 等ICML 2026 · 被引用 13 次
- Flow Matching for Multimodal DistributionsGaoxiang Luo, Frank Cole, Sihang Zhang, Yuxiang Wan 等CVPR 2026 · 被引用 1 次
- Evaluating the Representation Space of Diffusion Models via Self-Supervised PrinciplesXiao Li, Yixuan Jia, Zekai Zhang, Xiang Li 等ICML 2026
- Bridging RGB and RAW: Single-step Deterministic Flow with Homogeneous Representation AlignmentDiedong Feng, Peiyi Zeng, Zhen Liu, Zhongyang Li 等ICML 2026
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- U-REPA: Aligning Diffusion U-Nets to ViTsYuchuan Tian, Hanting Chen, Mengyu Zheng, Yuchen Liang 等NeurIPS 2025 · 被引用 30 次
- Dual-Path Condition Alignment for Diffusion TransformersChanghao Peng, Yuqi Ye, Shuangjun Du, Wenxu Gao 等ICLR 2026
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You ThinkGe Wu, Shen Zhang, Ruijing Shi, Shanghua Gao 等NeurIPS 2025 · 被引用 102 次
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You ThinkSihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong 等ICLR 2025
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion TransformersXingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing 等ICCV 2025 · 被引用 15 次
