Enabling Real-time Sign Language Translation on Mobile Platforms with On-board Depth Cameras
Hyeonjung Park, Youngki Lee, JeongGil Ko
摘要
In this work we present SUGO, a depth video-based system for translating sign language to text using a smartphone's front camera. While exploiting depth-only videos offer benefits such as being less privacy-invasive compared to using RGB videos, it introduces new challenges which include dealing with low video resolutions and the sensors' sensitiveness towards user motion. We overcome these challenges by diversifying our sign language video dataset to be robust to various usage scenarios via data augmentation and design a set of schemes to emphasize human gestures from the input images for effective sign detection. The inference engine of SUGO is based on a 3-dimensional convolutional neural network (3DCNN) to classify a sequence of video frames as a pre-trained word. Furthermore, the overall operations are designed to be light-weight so that sign language translation takes place in real-time using only the resources available on a smartphone, with no help from cloud servers nor external sensing components. Specifically, to train and test SUGO, we collect sign language data from 20 individuals for 50 Korean Sign Language words, summing up to a dataset of ∼5,000 sign gestures and collect additional in-the-wild data to evaluate the performance of SUGO in real-world usage scenarios with different lighting conditions and daily activities. Comprehensively, our extensive evaluations show that SUGO can properly classify sign words with an accuracy of up to 91% and also suggest that the system is suitable (in terms of resource usage, latency, and environmental robustness) to enable a fully mobile solution for sign language translation. CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InDepth: Real-time Depth Inpainting for Mobile Augmented RealityYunfan Zhang, Tim Scargill, Ashutosh Vaishnav, Gopika Premsankar 等UbiComp 2022 · 被引用 27 次
- Fast Monte-Carlo Approximation of the Attention MechanismHyunjun Kim, JeongGil KoAAAI 2022 · 被引用 8 次
- NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local LearningDhananjay Saikumar, Blesson VargheseEuroSys 2024 · 被引用 2 次
它引用的顶会 Paper3
- DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented RealityRuofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso 等UIST 2020 · 被引用 145 次
- mmASL: Environment-Independent ASL Gesture Recognition Using 60 GHz Millimeter-wave SignalsPanneer Selvam Santhalingam, Al Amin Hosain, Ding Zhang, Parth H. Pathak 等UbiComp 2020 · 被引用 89 次
- HeartQuake: Accurate Low-Cost Non-Invasive ECG Monitoring Using Bed-Mounted GeophonesJaeyeon Park, Hyeon Cho, Rajesh Krishna Balan, JeongGil KoUbiComp 2020 · 被引用 36 次
相关 Paper
- Reconstructing Signing Avatars from Video Using Linguistic PriorsMaria-Paola Forte, Peter Kulits, Chun-Hao Huang, Vasileios Choutas 等CVPR 2023
- WearSign: Pushing the Limit of Sign Language Translation Using Inertial and EMG WearablesQian Zhang, JiaZhen Jing, Dong Wang, Run ZhaoUbiComp 2022 · 被引用 26 次
- Skeleton-Aware Neural Sign Language TranslationShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie 等ACM MM 2021 · 被引用 28 次
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 被引用 58 次
- SignRing: Continuous American Sign Language Recognition Using IMU Rings and Virtual IMU DataJiyang Li, Lin Huang, Siddharth Shah, Sean J. Jones 等UbiComp 2023 · 被引用 27 次
