PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
Zhijie Wang, Yuheng Huang, Da Song, Lei Ma, Tianyi Zhang
摘要
The recent advancements in Generative AI have significantly advanced the field of text-to-image generation. The state-of-the-art text-to-image model, Stable Diffusion, is now capable of synthesizing high-quality images with a strong sense of aesthetics. Crafting text prompts that align with the model’s interpretation and the user’s intent thus becomes crucial. However, prompting remains challenging for novice users due to the complexity of the stable diffusion model and the non-trivial efforts required for iteratively editing and refining the text prompts. To address these challenges, we propose PromptCharm, a mixed-initiative system that facilitates text-to-image creation through multi-modal prompt engineering and refinement. To assist novice users in prompting, PromptCharm first automatically refines and optimizes the user’s initial prompt. Furthermore, PromptCharm supports the user in exploring and selecting different image styles within a large database. To assist users in effectively refining their prompts and images, PromptCharm renders model explanations by visualizing the model’s attention values. If the user notices any unsatisfactory areas in the generated images, they can further refine the images through model attention adjustment or image inpainting within the rich feedback loop of PromptCharm. To evaluate the effectiveness and usability of PromptCharm, we conducted a controlled user study with 12 participants and an exploratory user study with another 12 participants. These two studies show that participants using PromptCharm were able to create images with higher quality and better aligned with the user’s expectations compared with using two variants of PromptCharm that lacked interaction or visualization support.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood DiscrepancyShengfang Zhai, Huanran Chen, Yinpeng Dong, Jiajun Li 等NeurIPS 2024 · 被引用 48 次
- AIdeation: Designing a Human-AI Collaborative Ideation System for Concept DesignersWen-Fan Wang, Chien-Ting Lu, Nil Ponsa Campanyà, Bing-Yu Chen 等CHI 2025 · 被引用 46 次
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 被引用 33 次
- Exploring Collaboration Patterns and Strategies in Human-AI Co-creation through the Lens of Agency: A Scoping Review of the Top-tier HCI LiteratureShuning Zhang, Hui Wang, Xin YiCSCW 2025 · 被引用 31 次
- AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose ToolsNathalie Riche, Anna Offenwanger, Frederic Gmeiner, David Brown 等CHI 2025 · 被引用 26 次
它引用的顶会 Paper37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 被引用 303 次
- PromptMagician: Interactive Prompt Engineering for Text-to-Image CreationYingchaojie Feng, Xingbo Wang, Kamkwai Wong, Sijia Wang 等IEEE VIS 2023 · 被引用 127 次
- Dynamic Prompt Optimizing for Text-to-Image GenerationWenyi Mo, Tianyu Zhang, Yalong Bai, Bing Su 等CVPR 2024 · 被引用 15 次
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore 等UIST 2023 · 被引用 179 次
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like InteractionsJohn Joon Young Chung, Eytan AdarUIST 2023 · 被引用 94 次
