mSilent: Towards General Corpus Silent Speech Recognition Using COTS mmWave Radar
Shang Zeng, Haoran Wan, Shuyu Shi, Wei Wang
摘要
Silent speech recognition (SSR) allows users to speak to the device without making a sound, avoiding being overheard or disturbing others. Compared to the video-based approach, wireless signal-based SSR can work when the user is wearing a mask and has fewer privacy concerns. However, previous wireless-based systems are still far from well-studied, e.g. they are only evaluated in corpus with highly limited size, making them only feasible for interaction with dozens of deterministic commands. In this paper, we present mSilent, a millimeter-wave (mmWave) based SSR system that can work in the general corpus containing thousands of daily conversation sentences. With the strong recognition capability, mSilent not only supports the more complex interaction with assistants, but also enables more general applications in daily life such as communication and input. To extract fine-grained articulatory features, we build a signal processing pipeline that uses a clustering-selection algorithm to separate articulatory gestures and generates a multi-scale detrended spectrogram (MSDS). To handle the complexity of the general corpus, we design an end-to-end deep neural network that consists of a multi-branch convolutional front-end and a Transformer-based sequence-to-sequence back-end. We collect a general corpus dataset of 1,000 daily conversation sentences that contains 21K samples of bi-modality data (mmWave and video). Our evaluation shows that mSilent achieves a 9.5% average word error rate (WER) at a distance of 1.5m, which is comparable to the performance of the state-of-the-art video-based approach. We also explore deploying mSilent in two typical scenarios of text entry and in-car assistant, and the less than 6% average WER demonstrates the potential of mSilent in general daily applications.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools;
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MELDER: The Design and Evaluation of a Real-time Silent Speech Recognizer for Mobile DevicesLaxmi Pandey, Ahmed Sabbir ArifCHI 2024 · 被引用 14 次
- Exploring Uni-manual Around Ear Off-Device Gestures for EarablesShaikh Shawon Arefin Shimon, Ali Neshati, Junwei Sun, Qiang Xu 等UbiComp 2024 · 被引用 4 次
它引用的顶会 Paper21
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 被引用 212 次
- Contactless seismocardiography via deep learning radarsUnsoo Ha, Salah Assana, Fadel AdibMobiCom 2020 · 被引用 198 次
- MoVi-Fi: motion-robust vital signs waveform recovery via deep interpreted RF sensingZhe Chen, Tianyue Zheng, Chao Cai, Jun LuoMobiCom 2021 · 被引用 190 次
- mmTrack: Passive Multi-Person Localization Using Commodity Millimeter Wave RadioChenshu Wu, Feng Zhang, Beibei Wang, K. J. Ray LiuINFOCOM 2020 · 被引用 140 次
相关 Paper
- We Can Hear You with mmWave Radar! An End-to-End Eavesdropping SystemDachao Han, Teng Huang, Han Ding, Cui Zhao 等UbiComp 2026 · 被引用 5 次
- mmASL: Environment-Independent ASL Gesture Recognition Using 60 GHz Millimeter-wave SignalsPanneer Selvam Santhalingam, Al Amin Hosain, Ding Zhang, Parth H. Pathak 等UbiComp 2020 · 被引用 89 次
- EarCommand: "Hearing" Your Silent Speech Commands In EarYincheng Jin, Yang Gao, Xuhai Xu, Seokmin Choi 等UbiComp 2022 · 被引用 37 次
- ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled ExpressionsXuefu Dong, Yifei Chen, Yuuki Nishiyama, Kaoru Sezaki 等CHI 2024 · 被引用 20 次
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang 等CHI 2023 · 被引用 53 次
