mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count Prior
Dachao Han, Teng Huang, Han Ding, Cui Zhao, Fei Wang, Ge Wang, Wei Xi
Abstract
Automatic speech enhancement and separation are critical for improving speech quality and intelligibility in noisy, multi-speaker environments, with applications in hearing aids, automatic speech recognition (ASR), and voice-controlled systems. However, most existing methods rely on task-specific models designed for a fixed number of speakers—such as one model for single-speaker enhancement and others for two- or three-speaker separation—requiring explicit prior knowledge of speaker count. This design paradigm limits flexibility and scalability in real-world scenarios where the number of speakers is often unknown and dynamic. To overcome this limitation, we introduce Ra-dioSEP, a unified mmWave-audio multimodal framework that performs both speech enhancement and separation without prior knowledge of speaker count. By leveraging mmWave radar sensing, RadioSEP detects, localizes, and profiles multiple speakers, enabling dynamic estimation of speaker count and extraction of speaker-specific physical features. These features are fused with noisy audio and processed by a single deep neural network that adapts to the current speaker configuration, generating clean speech streams accordingly. Extensive experiments show that RadioSEP consistently outperforms state-of-the-art methods in both speech enhancement and multi-speaker separation tasks, while offering significantly improved generalization and adaptability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 540c5f78-2ee0-48a5-9a75-c753d90d59bdRelated papers
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- Who Speaks What from Afar: Eavesdropping In-Person Conversations via mmWave SensingShaoying Wang, Hansong Zhou, Yukun Yuan, Xiaonan ZhangINFOCOM 2026
- Radio2Text: Streaming Speech Recognition Using mmWave Radio SignalsRunning Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. NgaiUbiComp 2023 · 22 citations
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- Knowing When to Quit: Probabilistic Early Exits for Speech Separation NetworksKenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen et al.ICLR 2026 · 1 citation
