Complexity Matters: Rethinking the Latent Space for Generative Modeling
Tianyang Hu, Fei Chen, Haonan Wang, Jiawei Li, Wenjia Wang, Jiacheng Sun, Zhenguo Li
摘要
In generative modeling, numerous successful approaches leverage a lowdimensional latent space, e.g., Stable Diffusion [68] models the latent space induced by an encoder and generates images through a paired decoder. Although the selection of the latent space is empirically pivotal, determining the optimal choice and the process of identifying it remain unclear. In this study, we aim to shed light on this under-explored topic by rethinking the latent space from the perspective of model complexity. Our investigation starts with the classic generative adversarial networks (GANs). Inspired by the GAN training objective, we propose a novel "distance" between the latent and data distributions, whose minimization coincides with that of the generator complexity. The minimizer of this distance is characterized as the optimal data-dependent latent that most effectively capitalizes on the generator's capacity. Then, we consider parameterizing such a latent distribution by an encoder network and propose a two-stage training strategy called Decoupled Autoencoder (DAE), where the encoder is only updated in the first stage with an auxiliary decoder and then frozen in the second stage while the actual decoder is being trained. DAE can improve the latent distribution and as a result, improve the generative performance. Our theoretical analyses are corroborated by comprehensive experiments on various models such as VQGAN [21] and Diffusion Transformer [60], where our modifications yield significant improvements in sample quality with decreased model complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Stabilize the Latent Space for Image Autoregressive Modeling: A Unified PerspectiveYongxin Zhu, Bocheng Li, Hang Zhang, Xin Li 等NeurIPS 2024 · 被引用 26 次
- Elucidating the design space of classifier-guided diffusion generationJiajun Ma, Tianyang Hu, Wenjia Wang, Jiacheng SunICLR 2024 · 被引用 24 次
- FerretNet: Efficient Synthetic Image Detection via Local Pixel DependenciesShuqiao Liang, Jian Liu, Renzhang Chen, Quanlong GuanNeurIPS 2025 · 被引用 16 次
- The Surprising Effectiveness of Skip-Tuning in Diffusion SamplingJiajun Ma, Shuchen Xue, Tianyang Hu, Wenjia Wang 等ICML 2024 · 被引用 16 次
- PocketSR: The Super-Resolution Expert in Your Pocket MobilesHaoze Sun, Linfeng Jiang, Fan Li, Renjing Pei 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpaceJunyu Chen, Dongyun Zou, Wenkun He, Junsong Chen 等ICCV 2025 · 被引用 3 次
- Information-theoretic Generalization Analysis for VQ-VAEs: A Role of Latent VariablesFutoshi Futami, Masahiro FujisawaNeurIPS 2025 · 被引用 1 次
- Adversarial Latent AutoencodersStanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco DorettoCVPR 2020
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic AnalysisQi Chen, Jierui Zhu, Florian ShkurtiICLR 2025
- Perceptual Generative AutoencodersZijun Zhang, Ruixiang Zhang, Zongpeng Li, Yoshua Bengio 等ICML 2020 · 被引用 31 次
