Sensing to Hear: Speech Enhancement for Mobile Devices Using Acoustic Signals
Qian Zhang, Dong Wang, Run Zhao, Yinggang Yu, Junjie Shen
Abstract
Voice interactions and voice messages on mobile phones are rapidly growing in popularity. However, the user experience of these services is still worse than desired in noisy environments, especially in multi-talker scenarios, where the phone can only provide low-quality voice recordings. Speech enhancement using only audio as the input remains a grand challenge in these scenarios. In this paper, we handle this with the help of the emerging acoustic sensing technology. The key insight is that the inaudible acoustic signals emitted by speakers of phones can capture the subtle lip movements when people speak. Instead of enabling lip reading for the classification of limited voice commands, we further unlock the potential of acoustic sensing and leverage the captured lip information to improve the voice recording quality. We propose WaveVoice, a joint audio-sensory deep learning method for end-to-end speech enhancement on mobile phones. The model of WaveVoice is structured as an encoder-decoder network, in which audio and acoustic sensing data are processed through two individual CNN branches, respectively, and then fused into a joint network to generate enhanced speech. In addition, to improve the performance on new users, a self-supervised learning methodology is developed to adapt the model to extract speaker-specific features. We construct a dataset to train and evaluate WaveVoice. We also perform online tests under various noisy conditions to show the applicability of our system in real-world scenarios. Experimental results show that WaveVoice can effectively reconstruct the target clean speech from the noisy audio signals, and yield notably superior performance compared with the audio-only encoder-decoder model and the state-of-the-art speech enhancement methods. Given its promising performance, we believe that WaveVoice has made a substantial contribution to the advancement of mobile voice input.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6a5501dd-6671-421f-b216-4e3170b86233Cited by top-tier papers5
- PowerPhone: Unleashing the Acoustic Sensing Capability of SmartphonesShirui Cao, Dong Li, Sunghoon Ivan Lee, Jie XiongMobiCom 2023 · 35 citations
- ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic SensingPragma Kar, Shyamvanshikumar Singh, Avijit Mandal, Samiran Chattopadhyay et al.CHI 2023 · 7 citations
- AdaStreamLite: Environment-adaptive Streaming Speech Recognition on Mobile DevicesYuheng Wei, Jie Xiong, Hui Liu, Yingtao Yu et al.UbiComp 2024 · 5 citations
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai et al.UbiComp 2025 · 1 citation
- Exploring Acoustic Reverse Nonlinearity Against Speech Forgery in Real-Time Voice ApplicationsMing Gao, Lingfeng Zhang, Yike Chen, Sifeng He et al.INFOCOM 2025 · 1 citation
Related papers
- SoundLip: Enabling Word and Sentence-level Lip Interaction for Smart DevicesQian Zhang, Dong Wang, Run Zhao, Yinggang YuUbiComp 2021 · 34 citations
- Sensing to Hear through Memory: Ultrasound Speech Enhancement without Real Ultrasound SignalsQian Zhang, Ke Liu, Dong WangUbiComp 2024 · 5 citations
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 5 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- WearSE: Enabling Streaming Speech Enhancement on Eyewear Using Acoustic SensingQian Zhang, Kaiyi Guo, Yifei Yang, Dong WangUbiComp 2025 · 7 citations
