mmMUSE: An mmWave-based Motion-resilient Universal Speech Enhancement System
Lingyu Wang, Kai Wang, Dequan Wang, You Zuo, Chenming He, Chengzhen Meng, Xiaoran Fan, Haojie Ren, Yanyong Zhang
Abstract
Speech enhancement improves the user interaction experience in voice-based smart systems. While microphone-based speech perception is limited by airborne noise, mmWave is immune to such interference. However, user-device motion hinders mmWave-based vocal extraction, dispersing vocal signals and introducing distortions. In this paper, we propose mmMUSE , an mmWave-based motion-resilient universal speech enhancement system that integrates mmWave and audio. To mitigate motion interference, we propose a two-stage method for robust vocal vibration extraction. Moreover, by proposing the Vocal-Noise-Ratio metric to assess the prominence of the vocal vibration, we enable real-time voice activity detection. We also design a complex-valued network that includes an attention-based fusion network for cross-modal complementing and a time-frequency masking network for correcting amplitude and phase of speech to isolate noises. Using datasets from 46 participants, mmMUSE outperforms the state-of-the-art speech enhancement models by 26% in SISDR and 34% in STOI on average. It also achieves SISDR improvements of 16.72 dB, 17.93 dB, 14.93 dB, and 18.95 dB in controlled environments involving intense noise, extensive motion, multiple speakers, and various obstructive materials, respectively. Finally, in real-world scenarios, including running, public spaces, and driving, mmMUSE achieves WER below 10%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1c44cf7e-2033-4983-a2f4-278009b501c0Cited by top-tier papers1
Ask how each one uses itRelated papers
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li et al.INFOCOM 2023 · 41 citations
- WearSE: Enabling Streaming Speech Enhancement on Eyewear Using Acoustic SensingQian Zhang, Kaiyi Guo, Yifei Yang, Dong WangUbiComp 2025 · 7 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- AmbiEar: mmWave Based Voice Recognition in NLoS ScenariosJia Zhang, Yinian Zhou, Rui Xi, Shuai Li et al.UbiComp 2022 · 28 citations
- Wavesdropper: Through-wall Word Detection of Human Speech via Commercial mmWave DevicesChao Wang, Feng Lin, Zhongjie Ba, Fan Zhang et al.UbiComp 2022 · 40 citations
