mSilent: Towards General Corpus Silent Speech Recognition Using COTS mmWave Radar
Shang Zeng, Haoran Wan, Shuyu Shi, Wei Wang
Abstract
Silent speech recognition (SSR) allows users to speak to the device without making a sound, avoiding being overheard or disturbing others. Compared to the video-based approach, wireless signal-based SSR can work when the user is wearing a mask and has fewer privacy concerns. However, previous wireless-based systems are still far from well-studied, e.g. they are only evaluated in corpus with highly limited size, making them only feasible for interaction with dozens of deterministic commands. In this paper, we present mSilent, a millimeter-wave (mmWave) based SSR system that can work in the general corpus containing thousands of daily conversation sentences. With the strong recognition capability, mSilent not only supports the more complex interaction with assistants, but also enables more general applications in daily life such as communication and input. To extract fine-grained articulatory features, we build a signal processing pipeline that uses a clustering-selection algorithm to separate articulatory gestures and generates a multi-scale detrended spectrogram (MSDS). To handle the complexity of the general corpus, we design an end-to-end deep neural network that consists of a multi-branch convolutional front-end and a Transformer-based sequence-to-sequence back-end. We collect a general corpus dataset of 1,000 daily conversation sentences that contains 21K samples of bi-modality data (mmWave and video). Our evaluation shows that mSilent achieves a 9.5% average word error rate (WER) at a distance of 1.5m, which is comparable to the performance of the state-of-the-art video-based approach. We also explore deploying mSilent in two typical scenarios of text entry and in-car assistant, and the less than 6% average WER demonstrates the potential of mSilent in general daily applications.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools;
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a624137c-9c6d-4c92-a725-ec4f1ab9936fCited by top-tier papers2
- MELDER: The Design and Evaluation of a Real-time Silent Speech Recognizer for Mobile DevicesLaxmi Pandey, Ahmed Sabbir ArifCHI 2024 · 14 citations
- Exploring Uni-manual Around Ear Off-Device Gestures for EarablesShaikh Shawon Arefin Shimon, Ali Neshati, Junwei Sun, Qiang Xu et al.UbiComp 2024 · 4 citations
Builds on21
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- Contactless seismocardiography via deep learning radarsUnsoo Ha, Salah Assana, Fadel AdibMobiCom 2020 · 198 citations
- MoVi-Fi: motion-robust vital signs waveform recovery via deep interpreted RF sensingZhe Chen, Tianyue Zheng, Chao Cai, Jun LuoMobiCom 2021 · 190 citations
- mmTrack: Passive Multi-Person Localization Using Commodity Millimeter Wave RadioChenshu Wu, Feng Zhang, Beibei Wang, K. J. Ray LiuINFOCOM 2020 · 140 citations
Related papers
- We Can Hear You with mmWave Radar! An End-to-End Eavesdropping SystemDachao Han, Teng Huang, Han Ding, Cui Zhao et al.UbiComp 2026 · 5 citations
- mmASL: Environment-Independent ASL Gesture Recognition Using 60 GHz Millimeter-wave SignalsPanneer Selvam Santhalingam, Al Amin Hosain, Ding Zhang, Parth H. Pathak et al.UbiComp 2020 · 89 citations
- EarCommand: "Hearing" Your Silent Speech Commands In EarYincheng Jin, Yang Gao, Xuhai Xu, Seokmin Choi et al.UbiComp 2022 · 37 citations
- ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled ExpressionsXuefu Dong, Yifei Chen, Yuuki Nishiyama, Kaoru Sezaki et al.CHI 2024 · 20 citations
- EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic SensingRuidong Zhang, Ke Li, Yihong Hao, Yufan Wang et al.CHI 2023 · 53 citations
