Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
Dian Xie, Shitong Shao, Lichen Bai, Zikai Zhou, Bojun Cheng, Shuo Yang, Jun Wu, Zeke Xie
Abstract
Classifier-free guidance (CFG) has helped diffusion models achieve great conditional generation in various fields. Recently, more diffusion guidance methods have emerged with improved generation quality and human preference. However, can these emerging diffusion guidance methods really achieve solid and significant improvements? In this paper, we rethink recent progress on diffusion guidance. Our work mainly consists of four contributions. First, we reveal a critical evaluation pitfall that common human preference models exhibit a strong bias towards large guidance scales. Simply increasing the CFG scale can easily improve quantitative evaluation scores due to strong semantic alignment, even if image quality is severely damaged (e.g., oversaturation and artifacts). Second, we introduce a novel guidance-aware evaluation (GA-Eval) framework that employs effective guidance scale calibration to enable fair comparison between current guidance methods and CFG by identifying the effects orthogonal and parallel to CFG effects. Third, motivated by the evaluation pitfall, we design Transcendent Diffusion Guidance (TDG) method that can significantly improve human preference scores in the conventional evaluation framework but actually does not work in practice. Fourth, in extensive experiments, we empirically evaluate recent eight diffusion guidance methods within the conventional evaluation framework and the proposed GA-Eval framework. Notably, simply increasing the CFG scales can compete with most studied diffusion guidance methods, while all methods suffer severely from winning rate degradation over standard CFG. Our work would strongly motivate the community to rethink the evaluation paradigm and future directions of this field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf5d073c-eba0-4418-aad3-f44ddd0e5a70Cited by top-tier papers3
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based PerspectiveRui Huang, Shitong Shao, Zikai Zhou, Pukun Zhao et al.CVPR 2026 · 7 citations
- FastLightGen: Fast and Light Video Generation with Fewer Steps and ParametersShitong Shao, Yufei Gu, Zeke XieCVPR 2026 · 4 citations
- CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You ThinkZening Sun, Zhengpeng Xie, Lichen Bai, Shitong Shao et al.CVPR 2026 · 3 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion ModelsFu-Yun Wang, Yunhao Shui, Jingtan Piao, Keqiang Sun et al.ICLR 2025
- No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion ModelsSeyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, Romann M. WeberICLR 2025
- Overshoot and Shrinkage in Classifier-Free Guidance: From Theory to PracticeKrunoslav Lehman Pavasovic, Jakob Verbeek, Giulio Biroli, Marc MezardICLR 2026
- Improving Classifier-Free Guidance in Masked Diffusion: Low-Dim Theoretical Insights with High-Dim ImpactKevin Rojas, Ye He, Chieh-Hsin Lai, Yuhta Takida et al.ICLR 2026 · 11 citations
- Rethinking the Spatial Inconsistency in Classifier-Free Diffusion GuidanceDazhong Shen, Guanglu Song, Zeyue Xue, Fu-Yun Wang et al.CVPR 2024 · 12 citations
