PromptEnhancer: Taming Your Rewriter for Text-to-Image Generation via Fine-Grained Reward
Linqing Wang, Zhiyong Xu, Ximing Xing, Yiji Cheng, Zhiyuan Zhao, Donghao Li, Tiankai Hang, Zhenxi Li, Jiale Tao, Qixun Wang, Ruihuang Li, Comi Chen
Abstract
Recent advances in text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in generating high-fidelity images. However, these models often struggle to faithfully render complex user prompts, particularly in aspects such as attribute binding, negation, and compositional relationships. To address this challenge, we introduce PromptEnhancer, a novel and universal prompt rewriting framework that enhances any pre-trained T2I model.Specifically, we adopt a multi-stage training pipeline to systematically boost the rewriter's understanding and rewriting performance. In the first stage, we conduct supervised fine-tuning (SFT) using CoT-enabled data to enable the rewriter to generate structured, chain-of-thought-style responses. In the second stage, we design a task-specific reward model—AlignEvaluator—to further align user prompts with fine-grained preferences through GRPO.The AlignEvaluator is trained to provide explicit and fine-grained feedback based on a systematic taxonomy derived from common T2I failure cases. By optimizing the rewriter to maximize the reward from AlignEvaluator, our framework learns to generate prompts that T2I models can interpret more precisely. Furthermore, we introduce a comprehensive human-aligned benchmark to facilitate future research in this direction. Extensive experiments demonstrate that PromptEnhancer significantly improves image-text alignment across a wide range of semantic and compositional challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd8d35f7-c250-4f84-a727-ddef7a5cb93aBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference OptimizationZhuohan Liu, Wujian Peng, Yitong Chen, Zuxuan WuCVPR 2026 · 2 citations
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
- VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image SynthesisShiyu Wu, Mingzhen Sun, Weining Wang, Yequan Wang et al.ICLR 2026 · 7 citations
- Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and EditingRunze He, YIJI CHENG, Tiankai Hang, Zhimin Li et al.CVPR 2026 · 6 citations
- Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM EncodersSiqi Kou, Jiachun Jin, Zetong Zhou, YE MA et al.ICML 2026 · 13 citations
