Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals
Running Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. Ngai
Abstract
Millimeter wave (mmWave) based speech recognition provides more possibility for audio-related applications, such as conference speech transcription and eavesdropping. However, considering the practicality in real scenarios, latency and recognizable vocabulary size are two critical factors that cannot be overlooked. In this paper, we propose Radio2Text, the first mmWave-based system for streaming automatic speech recognition (ASR) with a vocabulary size exceeding 13,000 words. Radio2Text is based on a tailored streaming Transformer that is capable of effectively learning representations of speech-related features, paving the way for streaming ASR with a large vocabulary. To alleviate the deficiency of streaming networks unable to access entire future inputs, we propose the Guidance Initialization that facilitates the transfer of feature knowledge related to the global context from the non-streaming Transformer to the tailored streaming Transformer through weight inheritance. Further, we propose a cross-modal structure based on knowledge distillation (KD), named cross-modal KD, to mitigate the negative effect of low quality mmWave signals on recognition performance. In the cross-modal KD, the audio streaming Transformer provides feature and response guidance that inherit fruitful and accurate speech information to supervise the training of the tailored radio streaming Transformer. The experimental results show that our Radio2Text can achieve a character error rate of 5.7% and a word error rate of 9.4% for the recognition of a vocabulary consisting of over 13,000 words.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87f773e6-a446-4571-88f7-4025e42bc0c0Cited by top-tier papers6
- Continual Learning with Strategic Selection and Forgetting for Network Intrusion DetectionXinchen Zhang, Running Zhao, Zhihan Jiang, Handi Chen et al.INFOCOM 2025 · 26 citations
- Internal Cross-layer Gradients for Extending Homogeneity to Heterogeneity in Federated LearningYun-Hin Chan, Rui Zhou, Running Zhao, Zhihan Jiang et al.ICLR 2024 · 14 citations
- mmSpyVR: Exploiting mmWave Radar for Penetrating Obstacles to Uncover Privacy Vulnerability of Virtual RealityLuoyu Mei, Ruofeng Liu, Zhimeng Yin, Qingchuan Zhao et al.UbiComp 2025 · 13 citations
- RadEar: A Self-Supervised RF Backscatter System for Voice Eavesdropping and SeparationQijun Wang, Peihao Yan, Chunqi Qian, Huacheng ZengINFOCOM 2026 · 2 citations
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai et al.UbiComp 2025 · 1 citation
Builds on19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- How much Position Information Do Convolutional Neural Networks Encode?Md. Amirul Islam, Sen Jia, Neil D. B. BruceICLR 2020 · 392 citations
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
- mmVib: micrometer-level vibration measurement with mmwave radarChengkun Jiang, Junchen Guo, Yuan He, Meng Jin et al.MobiCom 2020 · 154 citations
- Towards Efficient 3D Object Detection with Knowledge DistillationJihan Yang, Shaoshuai Shi, Runyu Ding, Zhe Wang et al.NeurIPS 2022 · 76 citations
Related papers
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context ModelingJiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu et al.ICLR 2021 · 80 citations
- Heuristic-free Knowledge Distillation for Streaming ASR via Multi-modal TrainingJi Won YoonAAAI 2025
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao et al.INFOCOM 2026 · 1 citation
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao et al.NeurIPS 2025 · 6 citations
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar et al.ICML 2023 · 12 citations
