Sensing to Hear: Speech Enhancement for Mobile Devices Using Acoustic Signals
Qian Zhang, Dong Wang, Run Zhao, Yinggang Yu, Junjie Shen
摘要
Voice interactions and voice messages on mobile phones are rapidly growing in popularity. However, the user experience of these services is still worse than desired in noisy environments, especially in multi-talker scenarios, where the phone can only provide low-quality voice recordings. Speech enhancement using only audio as the input remains a grand challenge in these scenarios. In this paper, we handle this with the help of the emerging acoustic sensing technology. The key insight is that the inaudible acoustic signals emitted by speakers of phones can capture the subtle lip movements when people speak. Instead of enabling lip reading for the classification of limited voice commands, we further unlock the potential of acoustic sensing and leverage the captured lip information to improve the voice recording quality. We propose WaveVoice, a joint audio-sensory deep learning method for end-to-end speech enhancement on mobile phones. The model of WaveVoice is structured as an encoder-decoder network, in which audio and acoustic sensing data are processed through two individual CNN branches, respectively, and then fused into a joint network to generate enhanced speech. In addition, to improve the performance on new users, a self-supervised learning methodology is developed to adapt the model to extract speaker-specific features. We construct a dataset to train and evaluate WaveVoice. We also perform online tests under various noisy conditions to show the applicability of our system in real-world scenarios. Experimental results show that WaveVoice can effectively reconstruct the target clean speech from the noisy audio signals, and yield notably superior performance compared with the audio-only encoder-decoder model and the state-of-the-art speech enhancement methods. Given its promising performance, we believe that WaveVoice has made a substantial contribution to the advancement of mobile voice input.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- PowerPhone: Unleashing the Acoustic Sensing Capability of SmartphonesShirui Cao, Dong Li, Sunghoon Ivan Lee, Jie XiongMobiCom 2023 · 被引用 35 次
- ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic SensingPragma Kar, Shyamvanshikumar Singh, Avijit Mandal, Samiran Chattopadhyay 等CHI 2023 · 被引用 7 次
- AdaStreamLite: Environment-adaptive Streaming Speech Recognition on Mobile DevicesYuheng Wei, Jie Xiong, Hui Liu, Yingtao Yu 等UbiComp 2024 · 被引用 5 次
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai 等UbiComp 2025 · 被引用 1 次
- Exploring Acoustic Reverse Nonlinearity Against Speech Forgery in Real-Time Voice ApplicationsMing Gao, Lingfeng Zhang, Yike Chen, Sifeng He 等INFOCOM 2025 · 被引用 1 次
相关 Paper
- SoundLip: Enabling Word and Sentence-level Lip Interaction for Smart DevicesQian Zhang, Dong Wang, Run Zhao, Yinggang YuUbiComp 2021 · 被引用 34 次
- Sensing to Hear through Memory: Ultrasound Speech Enhancement without Real Ultrasound SignalsQian Zhang, Ke Liu, Dong WangUbiComp 2024 · 被引用 5 次
- ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear MicrophonesDong Ma, Ting Dang, Ming Ding, Rajesh BalanUbiComp 2024 · 被引用 5 次
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 被引用 69 次
- WearSE: Enabling Streaming Speech Enhancement on Eyewear Using Acoustic SensingQian Zhang, Kaiyi Guo, Yifei Yang, Dong WangUbiComp 2025 · 被引用 7 次
