TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Tiange Luo, Venkatesh Potluri, Anhong Guo
摘要
People who are blind or have low vision regularly use their hands to interact with the physical world to gain access to objects’ shape, size, weight, and texture. However, many rich visual features remain inaccessible through touch alone, making it difficult to distinguish similar objects, interpret visual affordances, and form a complete understanding of objects. In this work, we present TouchScribe, a system that augments hand-object interactions with automated live visual descriptions. We trained a custom egocentric hand interaction model to recognize both common gestures (e.g., grab to inspect, hold side-by-side to compare) and unique ones by blind people (e.g., point to explore color, or swipe to read available texts). Furthermore, TouchScribe provides real-time and adaptive feedback based on hand movement, from hand interaction states, to object labels, and to visual details. Our user study and technical evaluations demonstrate that TouchScribe can provide rich and useful descriptions to support object understanding. Finally, we discuss the implications of making live visual descriptions responsive to users’ physical reach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison 等CHI 2023 · 被引用 103 次
- Enabling Hand Gesture Customization on Wrist-Worn DevicesXuhai Xu, Jun Gong, Carolina Brum, Lilian Liang 等CHI 2022 · 被引用 82 次
- GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality ApplicationsTianyi Wang, Xun Qian, Fengming He, Xiyun Hu 等UIST 2021 · 被引用 75 次
- ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image CaptionsJaewook Lee, Jaylin Herskovitz, Yi-Hao Peng, Anhong GuoCHI 2022 · 被引用 55 次
相关 Paper
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 被引用 54 次
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry 等CHI 2024 · 被引用 37 次
- TacNote: Tactile and Audio Note-Taking for Non-Visual AccessWan-Chen Lee, Ching-Wen Hung, Chao-Hsien Ting, Peggy Chi 等UIST 2023 · 被引用 9 次
- Examining Visual Semantic Understanding in Blind and Low-Vision Technology UsersVenkatesh Potluri, Tadashi E. Grindeland, Jon E. Froehlich, Jennifer MankoffCHI 2021 · 被引用 42 次
- OmniScribe: Authoring Immersive Audio Descriptions for 360° VideosRuei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee 等UIST 2022 · 被引用 32 次
