Separating Voices from Multiple Sound Sources using 2D Microphone Array
Xinran Lu, Lei Xie, Fang Wang, Tao Gu, Chuyu Wang, Wei Wang, Sanglu Lu
Abstract
Voice assistant has been widely used for human-computer interaction and automatic meeting minutes. However, for multiple sound sources, the performance of speech recognition in voice assistant decreases dramatically. Therefore, it is crucial to separate multiple voices efficiently for an effective voice assistant application in multi-user scenarios. In this paper, we present a novel voice separation system using a 2D microphone array in multiple sound source scenarios. Specifically, we propose a spatial filtering-based method to iteratively estimate the Angle of Arrival (AoA) of each sound source and separate the voice signals with adaptive beamforming. We use BeamForming-based cross-Correlation (BF-Correlation) to accurately assess the performance of beamforming and automatically optimize the voice separation in the iterative framework. Different from cross-correlation, BF-Correlation further performs cross-correlation among the after-beamforming voice signals processed with each linear microphone array. In this way, the mutual interference from voice signals out of the specified direction can be effectively suppressed or mitigated via the spatial filtering technique. We implement a prototype system and evaluate its performance in real environments. Experimental results show that the average AoA error is 1.4 degree and the average ratio of automatic speech recognition accuracy is 90.2% in the presence of three sound sources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 639ff815-9fc2-4e23-9088-d1ccbee6b408Builds on3
- Voice localization using nearby wall reflectionsSheng Shen, Daguan Chen, Yu-Lin Wei, Zhijian Yang et al.MobiCom 2020 · 79 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- We Hear Your PACE: Passive Acoustic Localization of Multiple Walking PersonsChao Cai, Henglin Pu, Peng Wang, Zhe Chen et al.UbiComp 2021 · 59 citations
Related papers
- MAVL: Multiresolution Analysis of Voice LocalizationMei Wang, Wei Sun, Lili QiuNSDI 2021 · 46 citations
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li et al.INFOCOM 2023 · 41 citations
- Learning to Separate Voices by Spatial RegionsAlan Xu, Romit Roy ChoudhuryICML 2022 · 18 citations
- DeepEar: Sound Localization with Binaural MicrophonesQiang Yang, Yuanqing ZhengINFOCOM 2022 · 8 citations
