Opal: Multimodal Image Generation for News Illustration
Vivian Liu, Han Qiao, Lydia B. Chilton
摘要
Advances in multimodal AI have presented people with powerful ways to create images from text. Recent work has shown that text-to-image generations are able to represent a broad range of subjects and artistic styles. However, finding the right visual language for text prompts is difficult. In this paper, we address this challenge with Opal, a system that produces text-to-image generations for news illustration. Given an article, Opal guides users through a structured search for visual concepts and provides a pipeline allowing users to generate illustrations based on an article’s tone, keywords, and related artistic styles. Our evaluation shows that Opal efficiently generates diverse sets of news illustrations, visual assets, and concept ideas. Users with Opal generated two times more usable results than users without. We discuss how structured exploration can help users better understand the capabilities of human AI co-creative systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore 等UIST 2023 · 被引用 179 次
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li 等CHI 2024 · 被引用 143 次
- PromptMagician: Interactive Prompt Engineering for Text-to-Image CreationYingchaojie Feng, Xingbo Wang, Kamkwai Wong, Sijia Wang 等IEEE VIS 2023 · 被引用 127 次
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise ExpressionsYunlong Wang, Shuyuan Shen, Brian Y. LimCHI 2023 · 被引用 118 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsVivian Liu, Lydia B. ChiltonCHI 2022 · 被引用 586 次
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 被引用 465 次
相关 Paper
- DesignWeaver: Dimensional Scaffolding for Text-to-Image Product DesignSirui Tao, Ivan Liang, Cindy Peng, Zhiqing Wang 等CHI 2025 · 被引用 19 次
- PromptCharm: Text-to-Image Generation through Multi-modal Prompting and RefinementZhijie Wang, Yuheng Huang, Da Song, Lei Ma 等CHI 2024 · 被引用 111 次
- LLMs Behind the Scenes: Enabling Narrative Scene IllustrationMelissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun 等EMNLP 2025
- Is It AI or Is It Me? Understanding Users' Prompt Journey with Text-to-Image Generative AI ToolsAtefeh Mahdavi Goloujeh, Anne Sullivan, Brian MagerkoCHI 2024 · 被引用 88 次
- Discovering Divergent Representations Between Text-To-Image ModelsLisa Dunlap, Joseph E. Gonzalez, Trevor Darrell, Fabian Caba Heilbron 等ICCV 2025 · 被引用 2 次
