mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count Prior
Dachao Han, Teng Huang, Han Ding, Cui Zhao, Fei Wang, Ge Wang, Wei Xi
摘要
Automatic speech enhancement and separation are critical for improving speech quality and intelligibility in noisy, multi-speaker environments, with applications in hearing aids, automatic speech recognition (ASR), and voice-controlled systems. However, most existing methods rely on task-specific models designed for a fixed number of speakers—such as one model for single-speaker enhancement and others for two- or three-speaker separation—requiring explicit prior knowledge of speaker count. This design paradigm limits flexibility and scalability in real-world scenarios where the number of speakers is often unknown and dynamic. To overcome this limitation, we introduce Ra-dioSEP, a unified mmWave-audio multimodal framework that performs both speech enhancement and separation without prior knowledge of speaker count. By leveraging mmWave radar sensing, RadioSEP detects, localizes, and profiles multiple speakers, enabling dynamic estimation of speaker count and extraction of speaker-specific physical features. These features are fused with noisy audio and processed by a single deep neural network that adapts to the current speaker configuration, generating clean speech streams accordingly. Extensive experiments show that RadioSEP consistently outperforms state-of-the-art methods in both speech enhancement and multi-speaker separation tasks, while offering significantly improved generalization and adaptability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 被引用 69 次
- Who Speaks What from Afar: Eavesdropping In-Person Conversations via mmWave SensingShaoying Wang, Hansong Zhou, Yukun Yuan, Xiaonan ZhangINFOCOM 2026
- Radio2Text: Streaming Speech Recognition Using mmWave Radio SignalsRunning Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. NgaiUbiComp 2023 · 被引用 22 次
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- Knowing When to Quit: Probabilistic Early Exits for Speech Separation NetworksKenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen 等ICLR 2026 · 被引用 1 次
