GenAssist: Making Image Generation Accessible
Mina Huh, Yi-Hao Peng, Amy Pavel
摘要
Blind and low vision (BLV) creators use images to communicate with sighted audiences. However, creating or retrieving images is challenging for BLV creators as it is difficult to use authoring tools or assess image search results. Thus, creators limit the types of images they create or recruit sighted collaborators. While text-to-image generation models let creators generate high-fidelity images based on a text description (i.e. prompt), it is difficult to assess the content and quality of generated images. We present GenAssist, a system to make text-to-image generation accessible. Using our interface, creators can verify whether generated image candidates followed the prompt, access additional details in the image not specified in the prompt, and skim a summary of similarities and differences between image candidates. To power the interface, GenAssist uses a large language model to generate visual questions, vision-language models to extract answers, and a large language model to summarize the results. Our study with 12 BLV creators demonstrated that GenAssist enables and simplifies the process of image selection and generation, making visual authoring more accessible to all.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 被引用 54 次
- ContextCam: Bridging Context Awareness with Creative Human-AI Image Co-CreationXianzhe Fan, Zihan Wu, Chun Yu, Fenggui Rao 等CHI 2024 · 被引用 49 次
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry 等CHI 2024 · 被引用 37 次
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 被引用 26 次
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin 等CHI 2025 · 被引用 26 次
它引用的顶会 Paper9
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise ExpressionsYunlong Wang, Shuyuan Shen, Brian Y. LimCHI 2023 · 被引用 118 次
- VoxLens: Making Online Data Visualizations Accessible with an Interactive JavaScript Plug-InAther Sharif, Olivia H. Wang, Alida T. Muongchan, Katharina Reinecke 等CHI 2022 · 被引用 88 次
- Accessibility of High-Fidelity Prototyping ToolsJunchen Li, Garreth W. Tigwell, Kristen ShinoharaCHI 2021 · 被引用 42 次
- Generating Audio-Visual Slideshows from Text Articles Using Word ConcretenessMackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh AgrawalaCHI 2020 · 被引用 36 次
- Crosspower: Bridging Graphics and LinguisticsHaijun XiaUIST 2020 · 被引用 32 次
相关 Paper
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu 等CHI 2026 · 被引用 1 次
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi 等CHI 2025 · 被引用 23 次
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision CreatorsFranklin Mingzhe Li, Michael Xieyang Liu, Cynthia L. Bennett, Shaun K. KaneCHI 2026 · 被引用 2 次
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won 等CHI 2026 · 被引用 3 次
