mmMIC: Multi-modal Speech Recognition based on mmWave Radar
Long Fan, Lei Xie, Xinran Lu, Yi Li, Chuyu Wang, Sanglu Lu
摘要
With the proliferation of voice assistants, microphone-based speech recognition technology usually cannot achieve good performance in the situation of multiple sound sources and ambient noises. In this paper, we propose a novel mmWave-based solution to perform speech recognition to tackle the issues of multiple sound sources and ambient noises, by precisely extracting the multi-modal features from lip motion and vocal-cords vibration from the single channel of mmWave. We propose a difference-based method for feature extraction of lip motion to suppress the dynamic interference from body motion and head motion. We propose a speech detection method based on cross-validation of lip motion and vocal-cords vibration so as to avoid wasting computing resources on nonspeaking activities. We propose a multi-modal fusion framework for speech recognition by fusing the signal features from lip motion and vocal-cords vibration with the attention mechanism. We implemented a prototype system and evaluated the performance in real test-beds. Experiment results show that the average speech recognition accuracy is 92.8% in realistic environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Radio2Text: Streaming Speech Recognition Using mmWave Radio SignalsRunning Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. NgaiUbiComp 2023 · 被引用 22 次
- Beamforming for Sensing: Hybrid Beamforming based on Transmitter-Receiver Collaboration for Millimeter-Wave SensingLong Fan, Lei Xie, Wenhui Zhou, Chuyu Wang 等UbiComp 2024 · 被引用 13 次
- RF-Parrot: Wireless Eavesdropping on Wired AudioYanni Yang, Genglin Wang, Zhenlin An, Guoming Zhang 等INFOCOM 2024 · 被引用 11 次
- Talk2Radar: Talking to mmWave Radars via Smartphone SpeakerKaiyan Cui, Leming Shen, Yuanqing Zheng, Fu Xiao 等INFOCOM 2024 · 被引用 8 次
- Facial Landmark Detection Based on High Precision Spatial Sampling via Millimeter-wave RadarYi Li, Chuyu Wang, Lei Xie, Qiancheng Jin 等UbiComp 2025 · 被引用 7 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- An Empirical Study of Training End-to-End Vision-and-Language TransformersZi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang 等CVPR 2022 · 被引用 313 次
- mmVib: micrometer-level vibration measurement with mmwave radarChengkun Jiang, Junchen Guo, Yuan He, Meng Jin 等MobiCom 2020 · 被引用 154 次
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 被引用 69 次
- mmPhone: Acoustic Eavesdropping on Loudspeakers via mmWave-characterized Piezoelectric EffectChao Wang, Feng Lin, Tiantian Liu, Ziwei Liu 等INFOCOM 2022 · 被引用 46 次
相关 Paper
- mmMUSE: An mmWave-based Motion-resilient Universal Speech Enhancement SystemLingyu Wang, Kai Wang, Dequan Wang, You Zuo 等UbiComp 2026 · 被引用 2 次
- WavoID: Robust and Secure Multi-modal User Identification via mmWave-voice MechanismTiantian Liu, Feng Lin, Chao Wang, Chenhan Xu 等UIST 2023 · 被引用 19 次
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 被引用 86 次
- AmbiEar: mmWave Based Voice Recognition in NLoS ScenariosJia Zhang, Yinian Zhou, Rui Xi, Shuai Li 等UbiComp 2022 · 被引用 28 次
