Entropy Rectifying Guidance for Diffusion and Flow Models
Tariq Berrada, Adriana Romero-Soriano, Michal Drozdzal, Jakob J. Verbeek, Karteek Alahari
摘要
Guidance techniques are commonly used in diffusion and flow models to improve image quality and input consistency for conditional generative tasks such as class-conditional and text-to-image generation. In particular, classifier-free guidance (CFG) is the most widely adopted guidance technique. It results, however, in trade-offs across quality, diversity and consistency: improving some at the expense of others. While recent work has shown that it is possible to disentangle these factors to some extent, such methods come with an overhead of requiring an additional (weaker) model, or require more forward passes per sampling step. In this paper, we propose Entropy Rectifying Guidance (ERG), a simple and effective guidance method based on inference-time changes in the attention mechanism of state-of-the-art diffusion transformer architectures, which allows for simultaneous improvements over image quality, diversity and prompt consistency. ERG is more general than CFG and similar guidance techniques, as it extends to unconditional sampling. We show that ERG results in significant improvements in various tasks, including text-to-image, class-conditional and unconditional image generation. We also show that ERG can be seamlessly combined with other recent guidance methods such as CADS and APG, further improving generation results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Inference-time Physics Alignment of Video Generative Models with Latent World ModelsJianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez 等CVPR 2026 · 被引用 32 次
- It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion ModelsAnne Harrington, A. Sophia Koepke, Shyamgopal Karthik, Trevor Darrell 等CVPR 2026 · 被引用 12 次
- Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image GenerationJohannes Schusterbauer, Ming Gui, Yusong Li, Pingchuan Ma 等CVPR 2026 · 被引用 4 次
- GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image GenerationYe Zhu, Kaleb Newman, Johannes Lutzeyer, Adriana Romero-Soriano 等ICML 2026 · 被引用 1 次
- GuidedBridge: Training-freely Improving Bridge Models with Prior GuidanceZehua Chen, Yucheng Yang, Binjie Yuan, Kaiwen Zheng 等ICML 2026
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion ModelsSeyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, Romann M. WeberICLR 2025
- Guiding a Diffusion Model by Swapping Its TokensWeijia Zhang, Yuehao Liu, Shanyan Guan, Wu Ran 等CVPR 2026 · 被引用 2 次
- Guiding Diffusion Models with Semantically Degraded ConditionsShilong Han, Yuming Zhang, Hongxia WangCVPR 2026 · 被引用 1 次
- Guiding a Diffusion Model with a Bad Version of ItselfTero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen 等NeurIPS 2024 · 被引用 338 次
- Improving Classifier-Free Guidance in Masked Diffusion: Low-Dim Theoretical Insights with High-Dim ImpactKevin Rojas, Ye He, Chieh-Hsin Lai, Yuhta Takida 等ICLR 2026 · 被引用 11 次
