Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
Yucheng Zhou, Hao Li, Jianbing Shen
Abstract
Recent studies have explored autoregressive models for image generation, with promising results, and have combined diffusion models with autoregressive frameworks to optimize image generation via diffusion losses. In this study, we present a theoretical analysis of diffusion and autoregressive models with diffusion loss, highlighting the latter's advantages. We present a theoretical comparison of conditional diffusion and autoregressive diffusion with diffusion loss, demonstrating that patch denoising optimization in autoregressive models effectively mitigates condition errors and leads to a stable condition distribution. Our analysis also reveals that autoregressive condition generation refines the condition, causing the condition error influence to decay exponentially. In addition, we introduce a novel condition refinement approach based on Optimal Transport (OT) theory to address ``condition inconsistency''. We theoretically demonstrate that formulating condition refinement as a Wasserstein Gradient Flow ensures convergence toward the ideal condition distribution, effectively mitigating condition inconsistency. Experiments demonstrate the superiority of our method over diffusion and autoregressive models with diffusion loss methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45aa3f97-3a80-4a79-a8f4-2fb3551042deCited by top-tier papers1
Ask how each one uses itBuilds on17
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Text Diffusion with Reinforced ConditioningYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang et al.AAAI 2024 · 2 citations
- Generative Modeling with Optimal Transport MapsLitu Rout, Alexander Korotin, Evgeny BurnaevICLR 2022 · 92 citations
- Bidirectional Temporal Diffusion Model for Temporally Consistent Human AnimationTserendorj Adiya, Jae Shin Yoon, Jungeun Lee, Sanghun Kim et al.ICLR 2024 · 2 citations
- Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsFei Shen, Hu Ye, Sibo Liu, Jun Zhang et al.AAAI 2025 · 74 citations
- Wasserstein-Aware Transfer: Class-Level Alignment for Robust Diffusion Model AdaptationZixian Huang, Chuan-Xian RenAAAI 2026
