WorldScribe: Towards Context-Aware Live Visual Descriptions
Ruei-Che Chang, Yuxuan Liu, Anhong Guo
Abstract
Automated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-standing challenge in accessibility. In this work, we develop WorldScribe, a system that generates automated live real-world visual descriptions that are customizable and adaptive to users’ contexts: (i) WorldScribe’s descriptions are tailored to users’ intents and prioritized based on semantic relevance. (ii) WorldScribe is adaptive to visual contexts, e.g., providing consecutively succinct descriptions for dynamic scenes, while presenting longer and detailed ones for stable settings. (iii) WorldScribe is adaptive to sound contexts, e.g., increasing volume in noisy environments, or pausing when conversations start. Powered by a suite of vision, language, and sound recognition models, WorldScribe introduces a description generation pipeline that balances the tradeoffs between their richness and latency to support real-time use. The design of WorldScribe is informed by prior work on providing visual descriptions and a formative study with blind participants. Our user study and subsequent pipeline evaluation show that WorldScribe can provide real-time and fairly accurate visual descriptions to facilitate environment understanding that is adaptive and customized to users’ contexts. Finally, we discuss the implications and further steps toward making live visual descriptions more context-aware and humanized.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94b3d02c-8a6e-4a87-a527-18ec8424ad58Cited by top-tier papers25
- WanderGuide: Indoor Map-less Robotic Guide for Exploration by Blind PeopleMasaki Kuribayashi, Kohei Uehara, Allan Wang, Shigeo Morishima et al.CHI 2025 · 31 citations
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart GlassesRunze Cai, Nuwan Janaka, Hyeongcheol Kim, Yang Chen et al.CHI 2025 · 26 citations
- DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio DiscussionsShuchang Xu, Xiaofu Jin, Huamin Qu, Yukang YanCHI 2025 · 26 citations
- Are We On Track? AI-Assisted Active and Passive Goal Reflection During MeetingsXinyue Chen, Lev Tankelevitch, Rishi Vanukuru, Ava Elizabeth Scott et al.CHI 2025 · 23 citations
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
Builds on22
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image DescriptionsAbigale Stangl, Meredith Ringel Morris, Danna GurariCHI 2020 · 136 citations
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low et al.CHI 2020 · 123 citations
Related papers
- TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual DescriptionsRuei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap et al.CHI 2026 · 2 citations
- OmniScribe: Authoring Immersive Audio Descriptions for 360° VideosRuei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee et al.UIST 2022 · 32 citations
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry et al.CHI 2024 · 37 citations
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 72 citations
- Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista LandscapesChitralekha Gupta, Jing Peng, Ashwin Ram, Shreyas Sridhar et al.CHI 2026 · 1 citation
