mmMIC: Multi-modal Speech Recognition based on mmWave Radar
Long Fan, Lei Xie, Xinran Lu, Yi Li, Chuyu Wang, Sanglu Lu
Abstract
With the proliferation of voice assistants, microphone-based speech recognition technology usually cannot achieve good performance in the situation of multiple sound sources and ambient noises. In this paper, we propose a novel mmWave-based solution to perform speech recognition to tackle the issues of multiple sound sources and ambient noises, by precisely extracting the multi-modal features from lip motion and vocal-cords vibration from the single channel of mmWave. We propose a difference-based method for feature extraction of lip motion to suppress the dynamic interference from body motion and head motion. We propose a speech detection method based on cross-validation of lip motion and vocal-cords vibration so as to avoid wasting computing resources on nonspeaking activities. We propose a multi-modal fusion framework for speech recognition by fusing the signal features from lip motion and vocal-cords vibration with the attention mechanism. We implemented a prototype system and evaluated the performance in real test-beds. Experiment results show that the average speech recognition accuracy is 92.8% in realistic environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a881027d-c216-462c-9b97-56f267d3ce3eCited by top-tier papers9
- Radio2Text: Streaming Speech Recognition Using mmWave Radio SignalsRunning Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. NgaiUbiComp 2023 · 22 citations
- Beamforming for Sensing: Hybrid Beamforming based on Transmitter-Receiver Collaboration for Millimeter-Wave SensingLong Fan, Lei Xie, Wenhui Zhou, Chuyu Wang et al.UbiComp 2024 · 13 citations
- RF-Parrot: Wireless Eavesdropping on Wired AudioYanni Yang, Genglin Wang, Zhenlin An, Guoming Zhang et al.INFOCOM 2024 · 11 citations
- Talk2Radar: Talking to mmWave Radars via Smartphone SpeakerKaiyan Cui, Leming Shen, Yuanqing Zheng, Fu Xiao et al.INFOCOM 2024 · 8 citations
- Facial Landmark Detection Based on High Precision Spatial Sampling via Millimeter-wave RadarYi Li, Chuyu Wang, Lei Xie, Qiancheng Jin et al.UbiComp 2025 · 7 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- An Empirical Study of Training End-to-End Vision-and-Language TransformersZi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang et al.CVPR 2022 · 313 citations
- mmVib: micrometer-level vibration measurement with mmwave radarChengkun Jiang, Junchen Guo, Yuan He, Meng Jin et al.MobiCom 2020 · 154 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- mmPhone: Acoustic Eavesdropping on Loudspeakers via mmWave-characterized Piezoelectric EffectChao Wang, Feng Lin, Tiantian Liu, Ziwei Liu et al.INFOCOM 2022 · 46 citations
Related papers
- mmMUSE: An mmWave-based Motion-resilient Universal Speech Enhancement SystemLingyu Wang, Kai Wang, Dequan Wang, You Zuo et al.UbiComp 2026 · 2 citations
- WavoID: Robust and Secure Multi-modal User Identification via mmWave-voice MechanismTiantian Liu, Feng Lin, Chao Wang, Chenhan Xu et al.UIST 2023 · 19 citations
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 86 citations
- AmbiEar: mmWave Based Voice Recognition in NLoS ScenariosJia Zhang, Yinian Zhou, Rui Xi, Shuai Li et al.UbiComp 2022 · 28 citations
