Latent Diffusion Model without Variational Autoencoder
Minglei Shi, Haolin Wang, Wenzhao Zheng, Ziyang Yuan, Xiaoshi Wu, Xintao Wang, Pengfei Wan, Jie Zhou, Jiwen Lu
Abstract
Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with Variational Autoencoders (VAEs). While effective for highfidelity synthesis, this VAE+Diffusion paradigm still suffers from limited training and inference efficiency, along with poor transferability to broader vision tasks. These issues stem from a key limitation of VAE latent spaces: the lack of clear semantic separation and strong discriminative structure. Our analysis confirms that these properties are not only crucial for perception and understanding tasks, but also equally essential for the stable and efficient training of latent diffusion models. Motivated by this insight, we introduce SVG-a novel latent diffusion model without variational autoencoders, which unleashes Self-supervised representations for Visual Generation. SVG constructs a feature space with clear semantic discriminability by leveraging frozen DINO features, while a lightweight residual branch captures fine-grained details for high-fidelity reconstruction. Diffusion models are trained directly on this semantically structured latent space to facilitate more efficient learning. As a result, SVG enables accelerated diffusion training, supports few-step sampling, and improves generative quality. Experimental results further show that SVG preserves the semantic and discriminative capabilities of the underlying self-supervised representations, providing a principled pathway toward task-general, high-quality visual representations. (e) Inference Steps vs. FID-50K (f) Training Epochs vs. FID-50K VAE Enc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30ac2682-295d-446d-baa6-261cd59da468Cited by top-tier papers25
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng et al.CVPR 2026 · 82 citations
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and GenerationZhengrong Yue, Haiyu Zhang, Xiangyu Zeng, Boyu Chen et al.ICLR 2026 · 25 citations
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and EditingShilong Zhang, He Zhang, Zhifei Zhang, Chongjian GE et al.ICML 2026 · 19 citations
- Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion ModelsTianci Bi, Xiaoyi Zhang, Yan Lu, Nanning ZhengCVPR 2026 · 17 citations
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent DiffusionYueming Pan, Ruoyu Feng, Qi Dai, Yuqi Wang et al.CVPR 2026 · 15 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris et al.NeurIPS 2025 · 47 citations
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 288 citations
- USP: Unified Self-Supervised Pretraining for Image Generation and UnderstandingXiangxiang Chu, Renda Li, Yong WangICCV 2025 · 3 citations
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderXiaojie Li, Yang Zhao, Ming Li, Yancheng Zhang et al.CVPR 2026
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion ModelsSamuel Lavoie, Michael Noukhovitch, Aaron C. CourvilleNeurIPS 2025 · 3 citations
