Lune

NeurIPS2023顶会

Visual Programming for Step-by-Step Text-to-Image Generation and Evaluation

Jaemin Cho, Abhay Zala, Mohit Bansal

出版方
2023年份
62被引次数
23顶会引用

摘要

As large language models have demonstrated impressive performance in many domains, recent works have adopted language models (LMs) as controllers of visual modules for vision-and-language tasks. While existing work focuses on equipping LMs with visual understanding, we propose two novel interpretable/explainable visual programming frameworks for text-to-image (T2I) generation and evaluation. First, we introduce VPGEN, an interpretable step-by-step T2I generation framework that decomposes T2I generation into three steps: object/count generation, layout generation, and image generation. We employ an LM to handle the first two steps (object/count generation and layout generation), by finetuning it on textlayout pairs. Our step-by-step T2I generation framework provides stronger spatial control than end-to-end models, the dominant approach for this task. Furthermore, we leverage the world knowledge of pretrained LMs, overcoming the limitation of previous layout-guided T2I works that can only handle predefined object classes. We demonstrate that our VPGEN has improved control in counts/spatial relations/scales of objects than state-of-the-art T2I generation models. Second, we introduce VPEVAL, an interpretable and explainable evaluation framework for T2I generation based on visual programming. Unlike previous T2I evaluations with a single scoring model that is accurate in some skills but unreliable in others, VPEVAL produces evaluation programs that invoke a set of visual modules that are experts in different skills, and also provides visual+textual explanations of the evaluation results. Our analysis shows that VPEVAL provides a more humancorrelated evaluation for skill-specific and open-ended prompts than widely used single model-based evaluation. We hope that our work encourages future progress on interpretable/explainable generation and evaluation for T2I models. (a) VPGen: Step-by-Step T2I Generation "two Pikachus on a table" (b) VPEval: Explainable T2I Evaluation "three dogs in the image" imgGen(getLayout(getObjCounts(prompt), prompt), prompt) Visual Program countEval(img, "dog", "==3") Visual Program >>> countEval(img, "dog", "==3") False Evaluation Program >>> prompt = "two Pikachus on a table" >>> obj2count = getObjCounts(prompt) # "pikachu":

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext adb8e19f-4a67-448e-8436-2eec0081d1a5

引用它的顶会 Paper23

问问它们各自怎么用它

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖