SoundLip: Enabling Word and Sentence-level Lip Interaction for Smart Devices
Qian Zhang, Dong Wang, Run Zhao, Yinggang Yu
Abstract
As a natural and convenient interaction modality, voice input has now become indispensable to smart devices (e.g. mobile phones and smart appliances). However, voice input is strongly constrained by surroundings and may raise privacy leakage in public areas. In this paper, we present SoundLip, an end-to-end interaction system enabling users to interact with smart devices via silent voice input. The key insight is to use inaudible acoustic signals to capture the lip movements of users when they issue commands. Previous works have considered lip reading as a naive classification task and thus can only recognize individual words. In contrast, our proposed system enables lip reading at both word and sentence levels, which are more suitable for daily-life use. We exploit the built-in speakers and microphones of smart devices to emit acoustic signals and listen to their reflections, respectively. In order to better abstract representations from multi-frequency and multi-modality acoustic signals, we elaborate a hierarchical convolutional neural network (HCNN) to serve as the front-end as well as recognize individual word commands. Then, for the sentence-level recognition, we exploit a multi-task encoder-decoder network to get around temporal segmentation and output sentences in an end-to-end way. We evaluate SoundLip on 20 individual words and 70 sentences from 12 participants. Our system achieves an accuracy of 91.2% at word-level and a word error rate of 7.1% at sentence-level in both user-independent and environment-independent settings. Given its innovative solution and promising performance, we believe that SoundLip has made a significant contribution to the advancement of silent voice input technology.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5f17dbcc-ddf2-440d-bc8f-7132ac9a0706Cited by top-tier papers2
- mSilent: Towards General Corpus Silent Speech Recognition Using COTS mmWave RadarShang Zeng, Haoran Wan, Shuyu Shi, Wei WangUbiComp 2023 · 34 citations
- EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave RadarXin Yao, Kecheng Huang, Yimin Chen, Jiawei Guo et al.USENIX Security 2025
Related papers
- Sensing to Hear: Speech Enhancement for Mobile Devices Using Acoustic SignalsQian Zhang, Dong Wang, Run Zhao, Yinggang Yu et al.UbiComp 2021 · 22 citations
- Lipwatch: Enabling Silent Speech Recognition on Smartwatches using Acoustic SensingQian Zhang, Yubin Lan, Kaiyi Guo, Dong WangUbiComp 2024 · 20 citations
- Endophasia: Utilizing Acoustic-Based Imaging for Issuing Contact-Free Silent Speech CommandsYongzhao Zhang, Wei-Hsiang Huang, Chih-Yun Yang, Wen-Ping Wang et al.UbiComp 2020 · 43 citations
- Watch Your Mouth: Silent Speech Recognition with Depth SensingXue Wang, Zixiong Su, Jun Rekimoto, Yang ZhangCHI 2024 · 19 citations
- EarCommand: "Hearing" Your Silent Speech Commands In EarYincheng Jin, Yang Gao, Xuhai Xu, Seokmin Choi et al.UbiComp 2022 · 37 citations
