Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
Subin Kim, Sangwoo Mo, Mamshad Nayeem Rizve, Yiran Xu, Difan Liu, Jinwoo Shin, Tobias Hinz
Abstract
Achieving precise alignment between user intent and generated visuals remains a central challenge in text-to-visual generation, as a single attempt often fails to produce the desired output. To handle this, prior approaches mainly scale the visual generation process (e.g., increasing sampling steps or seeds), but this quickly leads to a quality plateau. This limitation arises because prompt adaptation does not scale along with the increasing population of generated samples. To address this, we propose Prompt Redesign for Inferencetime Scaling, coined PRIS, a framework that adaptively revises the prompt during inference in response to the scaled visual generations. The core idea of PRIS is to review the generated visuals, identify recurring failure patterns across visuals, and redesign the prompt accordingly before regenerating the visuals with the revised prompt. To provide precise alignment feedback for prompt revision, we introduce a new verifier, element-level factual correction, which evaluates the alignment between prompt attributes and generated visuals at a fine-grained level, achieving more accurate and interpretable assessments than holistic measures. Extensive experiments on both text-to-image and text-to-video benchmarks demonstrate the effectiveness of our approach, including a 15% gain on VBench 2.0. These results highlight that jointly scaling prompts and visuals is key to fully leveraging scaling laws at inference-time. Visualizations are available at the website: https://subin-kim-cv.github.io/PRIS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06a165e7-68e9-4dd7-b1bf-663e697e116eBuilds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 303 citations
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan et al.NeurIPS 2025 · 284 citations
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
Related papers
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationJiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu et al.ICCV 2025 · 3 citations
- AlignVid: Taming Visual Dominance via Training-Free Attention Modulation in Text-guided Image-to-Video GenerationYexin Liu, Wenjie Shu, Zile Huang, Haoze Zheng et al.ICML 2026
- Progress by Pieces: Test-Time Scaling for Autoregressive Image GenerationJoonhyung Park, Hyeongwon Jang, Joowon Kim, Eunho YangCVPR 2026 · 2 citations
- RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image AlignmentLiyao Jiang, Ruichen Chen, Chao Gao, Di NiuCVPR 2026 · 7 citations
- The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationBingjie Gao, Xinyu Gao, Xiaoxue Wu, Yujie Zhou et al.CVPR 2025
