Toward Early Quality Assessment of Text-to-Image Diffusion Models
Huanlei Guo, Hongxin Wei, Bingyi Jing
Abstract
Recent text-to-image (T2I) diffusion and flow-matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate--then--select''mode: many seeds are sampled and only a few images are kept for use. However, this pipeline is highly resource-intensive since each candidate requires tens to hundreds of denoising steps, and evaluation metrics such as CLIPScore and ImageReward are post-hoc. In this work, we address this inefficiency by introducing Probe-Select, a plug-in module that enables efficient evaluation of image quality within the generation process. We observe that certain intermediate denoiser activations, even at early timesteps, encode a stable coarse structure, object layout and spatial arrangement--that strongly correlates with final image fidelity. Probe-Select exploits this property by predicting final quality scores directly from early activations, allowing unpromising seeds to be terminated early. Across diffusion and flow-matching backbones, our experiments show that early evaluation at only 20% of the trajectory accurately ranks candidate seeds and enables selective continuation. This strategy reduces sampling cost by over 60% while improving the quality of the retained images, demonstrating that early structural signals can effectively guide selective generation without altering the underlying generative model. Code is available at https://github.com/Guhuary/ProbeSelect.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38b3e130-2f25-45ae-a1bf-48adefc4cbefBuilds on19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Diffusion Probe: Generated Image Result Prediction Using CNN ProbesBukun Huang, Benlei Cui, Zhizeng Ye, Xuemei Dong et al.CVPR 2026 · 13 citations
- FlashEval: Towards Fast and Accurate Evaluation of Text-to-Image Diffusion Generative ModelsLin Zhao, Tianchen Zhao, Zinan Lin, Xuefei Ning et al.CVPR 2024 · 2 citations
- Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image GenerationZilyu Ye, Zhiyang Chen, Tiancheng Li, Zemin Huang et al.CVPR 2025
- TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion AccelerationHaowei Zhu, Tingxuan Huang, Xing Wang, Tianyu Zhao et al.CVPR 2026 · 3 citations
- Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image SetsDale Decatur, Thibault Groueix, Wang Yifan, Rana Hanocka et al.ICCV 2025
