Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
Jiatao Gu, Ying Shen, Shuangfei Zhai, Yizhe Zhang, Navdeep Jaitly, Joshua Susskind
Abstract
Diffusion models have emerged as a powerful tool for generating high-quality images from textual descriptions. Despite their successes, these models often exhibit limited diversity in the sampled images, particularly when sampling with a high classifier-free guidance weight. To address this issue, we present Kaleido, a novel approach that enhances the diversity of samples by incorporating autoregressive latent priors. Kaleido integrates an autoregressive language model that encodes the original caption and generates latent variables, serving as abstract and intermediary representations for guiding and facilitating the image generation process. In this paper, we explore a variety of discrete latent representations, including textual descriptions, detection bounding boxes, object blobs, and visual tokens. These representations diversify and enrich the input conditions to the diffusion models, enabling more diverse outputs. Our experimental results demonstrate that Kaleido effectively broadens the diversity of the generated image samples from a given textual description while maintaining high image quality. Furthermore, we show that Kaleido adheres closely to the guidance provided by the generated latent variables, demonstrating its capability to effectively control and direct the image generation process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 176c3506-b3d4-420e-bad9-9b628d4e74a3Cited by top-tier papers5
- STARFlow: Scaling Latent Normalizing Flows for High-resolution Image SynthesisJiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng et al.NeurIPS 2025 · 34 citations
- PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language ModelsRunze He, Bo Cheng, Yuhang Ma, Qingxiang Jia et al.ICCV 2025 · 1 citation
- The Power of Context: How Multimodality Improves Image Super-ResolutionKangfu Mei, Hossein Talebi, Mojtaba Ardakani, Vishal M. Patel et al.CVPR 2025
- Denoising Autoregressive Transformers for Scalable Text-to-Image GenerationJiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang et al.ICLR 2025
- Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image GenerationRaphael Tang, Xinyu Zhang, Lixinyu Xu, Yao Lu et al.EMNLP 2024
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Latent Diffusion for Language GenerationJustin Lovelace, Varsha Kishore, Chao Wan, Eliot Shekhtman et al.NeurIPS 2023 · 177 citations
- Diffusion Self-Guidance for Controllable Image GenerationDave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros et al.NeurIPS 2023 · 411 citations
- FastHybrid: Accelerating Hybrid Autoregressive Image Generation with Lookahead and Guided DecodingZhengguo Jiang, Fang Zhang, YongXiang Hua, Bocheng Li et al.CVPR 2026
- Self-Guided Diffusion ModelsVincent Tao Hu, David W. Zhang, Yuki M. Asano, Gertjan J. Burghouts et al.CVPR 2023
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion ModelsSamuel Lavoie, Michael Noukhovitch, Aaron C. CourvilleNeurIPS 2025 · 3 citations
