WorldScribe: Towards Context-Aware Live Visual Descriptions
Ruei-Che Chang, Yuxuan Liu, Anhong Guo
摘要
Automated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-standing challenge in accessibility. In this work, we develop WorldScribe, a system that generates automated live real-world visual descriptions that are customizable and adaptive to users’ contexts: (i) WorldScribe’s descriptions are tailored to users’ intents and prioritized based on semantic relevance. (ii) WorldScribe is adaptive to visual contexts, e.g., providing consecutively succinct descriptions for dynamic scenes, while presenting longer and detailed ones for stable settings. (iii) WorldScribe is adaptive to sound contexts, e.g., increasing volume in noisy environments, or pausing when conversations start. Powered by a suite of vision, language, and sound recognition models, WorldScribe introduces a description generation pipeline that balances the tradeoffs between their richness and latency to support real-time use. The design of WorldScribe is informed by prior work on providing visual descriptions and a formative study with blind participants. Our user study and subsequent pipeline evaluation show that WorldScribe can provide real-time and fairly accurate visual descriptions to facilitate environment understanding that is adaptive and customized to users’ contexts. Finally, we discuss the implications and further steps toward making live visual descriptions more context-aware and humanized.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- WanderGuide: Indoor Map-less Robotic Guide for Exploration by Blind PeopleMasaki Kuribayashi, Kohei Uehara, Allan Wang, Shigeo Morishima 等CHI 2025 · 被引用 31 次
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart GlassesRunze Cai, Nuwan Janaka, Hyeongcheol Kim, Yang Chen 等CHI 2025 · 被引用 26 次
- DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio DiscussionsShuchang Xu, Xiaofu Jin, Huamin Qu, Yukang YanCHI 2025 · 被引用 26 次
- Are We On Track? AI-Assisted Active and Passive Goal Reflection During MeetingsXinyue Chen, Lev Tankelevitch, Rishi Vanukuru, Ava Elizabeth Scott 等CHI 2025 · 被引用 23 次
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh 等UIST 2025 · 被引用 9 次
它引用的顶会 Paper22
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 被引用 457 次
- "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image DescriptionsAbigale Stangl, Meredith Ringel Morris, Danna GurariCHI 2020 · 被引用 136 次
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low 等CHI 2020 · 被引用 123 次
相关 Paper
- TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual DescriptionsRuei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap 等CHI 2026 · 被引用 2 次
- OmniScribe: Authoring Immersive Audio Descriptions for 360° VideosRuei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee 等UIST 2022 · 被引用 32 次
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry 等CHI 2024 · 被引用 37 次
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 被引用 72 次
- Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista LandscapesChitralekha Gupta, Jing Peng, Ashwin Ram, Shreyas Sridhar 等CHI 2026 · 被引用 1 次
