MELDER: The Design and Evaluation of a Real-time Silent Speech Recognizer for Mobile Devices
Laxmi Pandey, Ahmed Sabbir Arif
摘要
Silent speech is unaffected by ambient noise, increases accessibility, and enhances privacy and security. Yet current silent speech recognizers operate in a phrase-in/phrase-out manner, thus are slow, error prone, and impractical for mobile devices. We present MELDER, a Mobile Lip Reader that operates in real-time by splitting the input video into smaller temporal segments to process them individually. An experiment revealed that this substantially improves computation time, making it suitable for mobile devices. We further optimize the model for everyday use by exploiting the knowledge from a high-resource vocabulary using a transfer learning model. We then compare MELDER in both stationary and mobile settings with two state-of-the-art silent speech recognizers, where MELDER demonstrated superior overall performance. Finally, we compare two visual feedback methods of MELDER with the visual feedback method of Google Assistant. The outcomes shed light on how these proposed feedback methods influence users’ perceptions of the model’s performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 被引用 87 次
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang 等CHI 2023 · 被引用 53 次
- MuteIt: Jaw Motion Based Unvoiced Command Recognition Using EarableTanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 等UbiComp 2022 · 被引用 52 次
- Acceptability of Speech and Silent Speech Input Methods in Private and PublicLaxmi Pandey, Khalad Hasan, Ahmed Sabbir ArifCHI 2021 · 被引用 51 次
- C-Face: Continuously Reconstructing Facial Expressions by Deep Learning Contours of the Face with Ear-mounted Miniature CamerasTuochao Chen, Benjamin Steeper, Kinan Alsheikh, Songyun Tao 等UIST 2020 · 被引用 49 次
相关 Paper
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- LipLearner: Customizable Silent Speech Interactions on Mobile DevicesZixiong Su, Shitao Fang, Jun RekimotoCHI 2023 · 被引用 37 次
- LipType: A Silent Speech Recognizer Augmented with an Independent Repair ModelLaxmi Pandey, Ahmed Sabbir ArifCHI 2021 · 被引用 38 次
- Lipwatch: Enabling Silent Speech Recognition on Smartwatches using Acoustic SensingQian Zhang, Yubin Lan, Kaiyi Guo, Dong WangUbiComp 2024 · 被引用 20 次
