Improving the Diffusability of Autoencoders
Ivan Skorokhodov, Sharath Girish, Benran Hu, Willi Menapace, Yanyu Li, Rameen Abdal, Sergey Tulyakov, Aliaksandr Siarohin
摘要
Latent diffusion models have emerged as the leading approach for generating high-quality images and videos, utilizing compressed latent representations to reduce the computational burden of the diffusion process. While recent advancements have primarily focused on scaling diffusion backbones and improving autoencoder reconstruction quality, the interaction between these components has received comparatively less attention. In this work, we perform a spectral analysis of modern autoencoders and identify inordinate high-frequency components in their latent spaces, which are especially pronounced in the autoencoders with a large bottleneck channel size. We hypothesize that this high-frequency component interferes with the coarse-to-fine nature of the diffusion synthesis process and hinders the generation quality. To mitigate the issue, we propose scale equivariance: a simple regularization strategy that aligns latent and RGB spaces across frequencies by enforcing scale equivariance in the decoder. It requires minimal code changes and only up to 20K autoencoder fine-tuning steps, yet significantly improves generation quality, reducing FID by 19% for image generation on ImageNet-1K 256 2 and FVD by at least 44% for video generation on Kinetics-700 17 × 256 2 . The source code is available at https://github.com/ snap-research/diffusability .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 被引用 288 次
- Latent Diffusion Model without Variational AutoencoderMinglei Shi, Haolin Wang, Wenzhao Zheng, Ziyang Yuan 等ICLR 2026 · 被引用 85 次
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang 等CVPR 2026 · 被引用 59 次
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang 等CVPR 2026 · 被引用 46 次
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang 等ICLR 2026 · 被引用 36 次
它引用的顶会 Paper33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency PerspectiveBolin Lai, Xudong Wang, Saketh Rambhatla, James M. Rehg 等CVPR 2026 · 被引用 7 次
- EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image ModelingTheodoros Kouzelis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos KomodakisICML 2025
- Learnings from Scaling Visual Tokenizers for Reconstruction and GenerationPhilippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar 等ICML 2025
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpaceJunyu Chen, Dongyun Zou, Wenkun He, Junsong Chen 等ICCV 2025 · 被引用 3 次
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationFrançois Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe 等NeurIPS 2025 · 被引用 23 次
