MELDER: The Design and Evaluation of a Real-time Silent Speech Recognizer for Mobile Devices
Laxmi Pandey, Ahmed Sabbir Arif
Abstract
Silent speech is unaffected by ambient noise, increases accessibility, and enhances privacy and security. Yet current silent speech recognizers operate in a phrase-in/phrase-out manner, thus are slow, error prone, and impractical for mobile devices. We present MELDER, a Mobile Lip Reader that operates in real-time by splitting the input video into smaller temporal segments to process them individually. An experiment revealed that this substantially improves computation time, making it suitable for mobile devices. We further optimize the model for everyday use by exploiting the knowledge from a high-resource vocabulary using a transfer learning model. We then compare MELDER in both stationary and mobile settings with two state-of-the-art silent speech recognizers, where MELDER demonstrated superior overall performance. Finally, we compare two visual feedback methods of MELDER with the visual feedback method of Google Assistant. The outcomes shed light on how these proposed feedback methods influence users’ perceptions of the model’s performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e5d1853-e467-4f8d-b59c-5de4c7b072eaBuilds on12
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 87 citations
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang et al.CHI 2023 · 53 citations
- MuteIt: Jaw Motion Based Unvoiced Command Recognition Using EarableTanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen et al.UbiComp 2022 · 52 citations
- Acceptability of Speech and Silent Speech Input Methods in Private and PublicLaxmi Pandey, Khalad Hasan, Ahmed Sabbir ArifCHI 2021 · 51 citations
- C-Face: Continuously Reconstructing Facial Expressions by Deep Learning Contours of the Face with Ear-mounted Miniature CamerasTuochao Chen, Benjamin Steeper, Kinan Alsheikh, Songyun Tao et al.UIST 2020 · 49 citations
Related papers
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou et al.AAAI 2020 · 106 citations
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 104 citations
- LipLearner: Customizable Silent Speech Interactions on Mobile DevicesZixiong Su, Shitao Fang, Jun RekimotoCHI 2023 · 37 citations
- LipType: A Silent Speech Recognizer Augmented with an Independent Repair ModelLaxmi Pandey, Ahmed Sabbir ArifCHI 2021 · 38 citations
- Lipwatch: Enabling Silent Speech Recognition on Smartwatches using Acoustic SensingQian Zhang, Yubin Lan, Kaiyi Guo, Dong WangUbiComp 2024 · 20 citations
