Enabling Real-time Sign Language Translation on Mobile Platforms with On-board Depth Cameras
Hyeonjung Park, Youngki Lee, JeongGil Ko
Abstract
In this work we present SUGO, a depth video-based system for translating sign language to text using a smartphone's front camera. While exploiting depth-only videos offer benefits such as being less privacy-invasive compared to using RGB videos, it introduces new challenges which include dealing with low video resolutions and the sensors' sensitiveness towards user motion. We overcome these challenges by diversifying our sign language video dataset to be robust to various usage scenarios via data augmentation and design a set of schemes to emphasize human gestures from the input images for effective sign detection. The inference engine of SUGO is based on a 3-dimensional convolutional neural network (3DCNN) to classify a sequence of video frames as a pre-trained word. Furthermore, the overall operations are designed to be light-weight so that sign language translation takes place in real-time using only the resources available on a smartphone, with no help from cloud servers nor external sensing components. Specifically, to train and test SUGO, we collect sign language data from 20 individuals for 50 Korean Sign Language words, summing up to a dataset of ∼5,000 sign gestures and collect additional in-the-wild data to evaluate the performance of SUGO in real-world usage scenarios with different lighting conditions and daily activities. Comprehensively, our extensive evaluations show that SUGO can properly classify sign words with an accuracy of up to 91% and also suggest that the system is suitable (in terms of resource usage, latency, and environmental robustness) to enable a fully mobile solution for sign language translation. CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90da14a2-5dc3-46fb-9f0c-216f528f2854Cited by top-tier papers3
- InDepth: Real-time Depth Inpainting for Mobile Augmented RealityYunfan Zhang, Tim Scargill, Ashutosh Vaishnav, Gopika Premsankar et al.UbiComp 2022 · 27 citations
- Fast Monte-Carlo Approximation of the Attention MechanismHyunjun Kim, JeongGil KoAAAI 2022 · 8 citations
- NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local LearningDhananjay Saikumar, Blesson VargheseEuroSys 2024 · 2 citations
Builds on3
- DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented RealityRuofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso et al.UIST 2020 · 145 citations
- mmASL: Environment-Independent ASL Gesture Recognition Using 60 GHz Millimeter-wave SignalsPanneer Selvam Santhalingam, Al Amin Hosain, Ding Zhang, Parth H. Pathak et al.UbiComp 2020 · 89 citations
- HeartQuake: Accurate Low-Cost Non-Invasive ECG Monitoring Using Bed-Mounted GeophonesJaeyeon Park, Hyeon Cho, Rajesh Krishna Balan, JeongGil KoUbiComp 2020 · 36 citations
Related papers
- Reconstructing Signing Avatars from Video Using Linguistic PriorsMaria-Paola Forte, Peter Kulits, Chun-Hao Huang, Vasileios Choutas et al.CVPR 2023
- WearSign: Pushing the Limit of Sign Language Translation Using Inertial and EMG WearablesQian Zhang, JiaZhen Jing, Dong Wang, Run ZhaoUbiComp 2022 · 26 citations
- Skeleton-Aware Neural Sign Language TranslationShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie et al.ACM MM 2021 · 28 citations
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
- SignRing: Continuous American Sign Language Recognition Using IMU Rings and Virtual IMU DataJiyang Li, Lin Huang, Siddharth Shah, Sean J. Jones et al.UbiComp 2023 · 27 citations
