Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals
Running Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. Ngai
摘要
Millimeter wave (mmWave) based speech recognition provides more possibility for audio-related applications, such as conference speech transcription and eavesdropping. However, considering the practicality in real scenarios, latency and recognizable vocabulary size are two critical factors that cannot be overlooked. In this paper, we propose Radio2Text, the first mmWave-based system for streaming automatic speech recognition (ASR) with a vocabulary size exceeding 13,000 words. Radio2Text is based on a tailored streaming Transformer that is capable of effectively learning representations of speech-related features, paving the way for streaming ASR with a large vocabulary. To alleviate the deficiency of streaming networks unable to access entire future inputs, we propose the Guidance Initialization that facilitates the transfer of feature knowledge related to the global context from the non-streaming Transformer to the tailored streaming Transformer through weight inheritance. Further, we propose a cross-modal structure based on knowledge distillation (KD), named cross-modal KD, to mitigate the negative effect of low quality mmWave signals on recognition performance. In the cross-modal KD, the audio streaming Transformer provides feature and response guidance that inherit fruitful and accurate speech information to supervise the training of the tailored radio streaming Transformer. The experimental results show that our Radio2Text can achieve a character error rate of 5.7% and a word error rate of 9.4% for the recognition of a vocabulary consisting of over 13,000 words.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Continual Learning with Strategic Selection and Forgetting for Network Intrusion DetectionXinchen Zhang, Running Zhao, Zhihan Jiang, Handi Chen 等INFOCOM 2025 · 被引用 26 次
- Internal Cross-layer Gradients for Extending Homogeneity to Heterogeneity in Federated LearningYun-Hin Chan, Rui Zhou, Running Zhao, Zhihan Jiang 等ICLR 2024 · 被引用 14 次
- mmSpyVR: Exploiting mmWave Radar for Penetrating Obstacles to Uncover Privacy Vulnerability of Virtual RealityLuoyu Mei, Ruofeng Liu, Zhimeng Yin, Qingchuan Zhao 等UbiComp 2025 · 被引用 13 次
- RadEar: A Self-Supervised RF Backscatter System for Voice Eavesdropping and SeparationQijun Wang, Peihao Yan, Chunqi Qian, Huacheng ZengINFOCOM 2026 · 被引用 2 次
- USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal SynthesisLuca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai 等UbiComp 2025 · 被引用 1 次
它引用的顶会 Paper19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- How much Position Information Do Convolutional Neural Networks Encode?Md. Amirul Islam, Sen Jia, Neil D. B. BruceICLR 2020 · 被引用 392 次
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang 等CVPR 2022 · 被引用 228 次
- mmVib: micrometer-level vibration measurement with mmwave radarChengkun Jiang, Junchen Guo, Yuan He, Meng Jin 等MobiCom 2020 · 被引用 154 次
- Towards Efficient 3D Object Detection with Knowledge DistillationJihan Yang, Shaoshuai Shi, Runyu Ding, Zhe Wang 等NeurIPS 2022 · 被引用 76 次
相关 Paper
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context ModelingJiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu 等ICLR 2021 · 被引用 80 次
- Heuristic-free Knowledge Distillation for Streaming ASR via Multi-modal TrainingJi Won YoonAAAI 2025
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao 等INFOCOM 2026 · 被引用 1 次
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao 等NeurIPS 2025 · 被引用 6 次
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar 等ICML 2023 · 被引用 12 次
