GenAssist: Making Image Generation Accessible
Mina Huh, Yi-Hao Peng, Amy Pavel
Abstract
Blind and low vision (BLV) creators use images to communicate with sighted audiences. However, creating or retrieving images is challenging for BLV creators as it is difficult to use authoring tools or assess image search results. Thus, creators limit the types of images they create or recruit sighted collaborators. While text-to-image generation models let creators generate high-fidelity images based on a text description (i.e. prompt), it is difficult to assess the content and quality of generated images. We present GenAssist, a system to make text-to-image generation accessible. Using our interface, creators can verify whether generated image candidates followed the prompt, access additional details in the image not specified in the prompt, and skim a summary of similarities and differences between image candidates. To power the interface, GenAssist uses a large language model to generate visual questions, vision-language models to extract answers, and a large language model to summarize the results. Our study with 12 BLV creators demonstrated that GenAssist enables and simplifies the process of image selection and generation, making visual authoring more accessible to all.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cb39c89-e379-439c-9eb5-a6d60f89c2a5Cited by top-tier papers22
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 54 citations
- ContextCam: Bridging Context Awareness with Creative Human-AI Image Co-CreationXianzhe Fan, Zihan Wu, Chun Yu, Fenggui Rao et al.CHI 2024 · 49 citations
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry et al.CHI 2024 · 37 citations
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 26 citations
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin et al.CHI 2025 · 26 citations
Builds on9
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise ExpressionsYunlong Wang, Shuyuan Shen, Brian Y. LimCHI 2023 · 118 citations
- VoxLens: Making Online Data Visualizations Accessible with an Interactive JavaScript Plug-InAther Sharif, Olivia H. Wang, Alida T. Muongchan, Katharina Reinecke et al.CHI 2022 · 88 citations
- Accessibility of High-Fidelity Prototyping ToolsJunchen Li, Garreth W. Tigwell, Kristen ShinoharaCHI 2021 · 42 citations
- Generating Audio-Visual Slideshows from Text Articles Using Word ConcretenessMackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh AgrawalaCHI 2020 · 36 citations
- Crosspower: Bridging Graphics and LinguisticsHaijun XiaUIST 2020 · 32 citations
Related papers
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu et al.CHI 2026 · 1 citation
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi et al.CHI 2025 · 23 citations
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision CreatorsFranklin Mingzhe Li, Michael Xieyang Liu, Cynthia L. Bennett, Shaun K. KaneCHI 2026 · 2 citations
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won et al.CHI 2026 · 3 citations
