Prompt Refinement with Image Pivot for Text-to-Image Generation
Jingtao Zhan, Qingyao Ai, Yiqun Liu, Yingwei Pan, Ting Yao, Jiaxin Mao, Shaoping Ma, Tao Mei
Abstract
For text-to-image generation, automatically refining user-provided natural language prompts into the keyword-enriched prompts favored by systems is essential for the user experience. Such a prompt refinement process is analogous to translating the prompt from "user languages" into "system languages". However, the scarcity of such parallel corpora makes it difficult to train a prompt refinement model. Inspired by zero-shot machine translation techniques, we introduce Prompt Refinement with Image Pivot (PRIP). PRIP innovatively uses the latent representation of a user-preferred image as an intermediary "pivot" between the user and system languages. It decomposes the refinement process into two data-rich tasks: inferring representations of user-preferred images from user languages and subsequently translating image representations into system languages. Thus, it can leverage abundant data for training. Extensive experiments show that PRIP substantially outperforms a wide range of baselines and effectively transfers to unseen systems in a zero-shot manner 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15685d0f-37d2-48d8-a718-c9cf9c04990bCited by top-tier papers5
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi et al.ACL 2025 · 13 citations
- Show, Don't Tell: Morphing Latent Reasoning into Image GenerationHarold Haodong Chen, Xinxiang Yin, Wenjie Shu, Hongfei (Faye) Zhang et al.ICML 2026 · 7 citations
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang et al.ICCV 2025 · 3 citations
- The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationBingjie Gao, Xinyu Gao, Xiaoxue Wu, Yujie Zhou et al.CVPR 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- MPPR: Memory-Prior-based Prompt Refinement in Continuous Space for Advanced Text-to-Image GenerationZhibing Zhang, Jiantao Lin, Cangqi Zhou, Rui XiaACM MM 2025
- RAt: Injecting Implicit Bias for Text-To-Image Prompt Refinement ModelsZiyi Kou, Shichao Pei, Meng Jiang, Xiangliang ZhangEMNLP 2024 · 1 citation
- LANIT: Language-Driven Image-to-Image Translation for Unlabeled DataJihye Park, Sunwoo Kim, Soohyun Kim, Seokju Cho et al.CVPR 2023
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image SynthesisNailei Hei, Qianyu Guo, Zihao Wang, Yan Wang et al.AAAI 2024 · 11 citations
- TIPO: Text to Image with Text Pre-sampling for Prompt OptimizationShih-Ying Yeh, YI LI, SangHyun Park, Giyeong Oh et al.ICLR 2026 · 2 citations
