Complexity Matters: Rethinking the Latent Space for Generative Modeling
Tianyang Hu, Fei Chen, Haonan Wang, Jiawei Li, Wenjia Wang, Jiacheng Sun, Zhenguo Li
Abstract
In generative modeling, numerous successful approaches leverage a lowdimensional latent space, e.g., Stable Diffusion [68] models the latent space induced by an encoder and generates images through a paired decoder. Although the selection of the latent space is empirically pivotal, determining the optimal choice and the process of identifying it remain unclear. In this study, we aim to shed light on this under-explored topic by rethinking the latent space from the perspective of model complexity. Our investigation starts with the classic generative adversarial networks (GANs). Inspired by the GAN training objective, we propose a novel "distance" between the latent and data distributions, whose minimization coincides with that of the generator complexity. The minimizer of this distance is characterized as the optimal data-dependent latent that most effectively capitalizes on the generator's capacity. Then, we consider parameterizing such a latent distribution by an encoder network and propose a two-stage training strategy called Decoupled Autoencoder (DAE), where the encoder is only updated in the first stage with an auxiliary decoder and then frozen in the second stage while the actual decoder is being trained. DAE can improve the latent distribution and as a result, improve the generative performance. Our theoretical analyses are corroborated by comprehensive experiments on various models such as VQGAN [21] and Diffusion Transformer [60], where our modifications yield significant improvements in sample quality with decreased model complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dbdc3ad-5be9-4ff9-a826-13d68a185319Cited by top-tier papers10
- Stabilize the Latent Space for Image Autoregressive Modeling: A Unified PerspectiveYongxin Zhu, Bocheng Li, Hang Zhang, Xin Li et al.NeurIPS 2024 · 26 citations
- Elucidating the design space of classifier-guided diffusion generationJiajun Ma, Tianyang Hu, Wenjia Wang, Jiacheng SunICLR 2024 · 24 citations
- FerretNet: Efficient Synthetic Image Detection via Local Pixel DependenciesShuqiao Liang, Jian Liu, Renzhang Chen, Quanlong GuanNeurIPS 2025 · 16 citations
- The Surprising Effectiveness of Skip-Tuning in Diffusion SamplingJiajun Ma, Shuchen Xue, Tianyang Hu, Wenjia Wang et al.ICML 2024 · 16 citations
- PocketSR: The Super-Resolution Expert in Your Pocket MobilesHaoze Sun, Linfeng Jiang, Fan Li, Renjing Pei et al.NeurIPS 2025 · 8 citations
Builds on30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpaceJunyu Chen, Dongyun Zou, Wenkun He, Junsong Chen et al.ICCV 2025 · 3 citations
- Information-theoretic Generalization Analysis for VQ-VAEs: A Role of Latent VariablesFutoshi Futami, Masahiro FujisawaNeurIPS 2025 · 1 citation
- Adversarial Latent AutoencodersStanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco DorettoCVPR 2020
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic AnalysisQi Chen, Jierui Zhu, Florian ShkurtiICLR 2025
- Perceptual Generative AutoencodersZijun Zhang, Ruixiang Zhang, Zongpeng Li, Yoshua Bengio et al.ICML 2020 · 31 citations
