TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Tiange Luo, Venkatesh Potluri, Anhong Guo
Abstract
People who are blind or have low vision regularly use their hands to interact with the physical world to gain access to objects’ shape, size, weight, and texture. However, many rich visual features remain inaccessible through touch alone, making it difficult to distinguish similar objects, interpret visual affordances, and form a complete understanding of objects. In this work, we present TouchScribe, a system that augments hand-object interactions with automated live visual descriptions. We trained a custom egocentric hand interaction model to recognize both common gestures (e.g., grab to inspect, hold side-by-side to compare) and unique ones by blind people (e.g., point to explore color, or swipe to read available texts). Furthermore, TouchScribe provides real-time and adaptive feedback based on hand movement, from hand interaction states, to object labels, and to visual details. Our user study and technical evaluations demonstrate that TouchScribe can provide rich and useful descriptions to support object understanding. Finally, we discuss the implications of making live visual descriptions responsive to users’ physical reach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 419cf8fb-0647-44a3-80d4-d1f636783b1cBuilds on20
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison et al.CHI 2023 · 103 citations
- Enabling Hand Gesture Customization on Wrist-Worn DevicesXuhai Xu, Jun Gong, Carolina Brum, Lilian Liang et al.CHI 2022 · 82 citations
- GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality ApplicationsTianyi Wang, Xun Qian, Fengming He, Xiyun Hu et al.UIST 2021 · 75 citations
- ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image CaptionsJaewook Lee, Jaylin Herskovitz, Yi-Hao Peng, Anhong GuoCHI 2022 · 55 citations
Related papers
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 54 citations
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry et al.CHI 2024 · 37 citations
- TacNote: Tactile and Audio Note-Taking for Non-Visual AccessWan-Chen Lee, Ching-Wen Hung, Chao-Hsien Ting, Peggy Chi et al.UIST 2023 · 9 citations
- Examining Visual Semantic Understanding in Blind and Low-Vision Technology UsersVenkatesh Potluri, Tadashi E. Grindeland, Jon E. Froehlich, Jennifer MankoffCHI 2021 · 42 citations
- OmniScribe: Authoring Immersive Audio Descriptions for 360° VideosRuei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee et al.UIST 2022 · 32 citations
