The Cone of Silence: Speech Separation by Localization
Teerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-Shlizerman
Abstract
Given a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the waveform domain, which isolates sources within an angular region , given an angle of interest and angular window size . By exponentially decreasing , we can perform a binary search to localize and separate all sources in logarithmic time. Our algorithm allows for an arbitrary number of potentially moving speakers at test time, including more speakers than seen during training. Experiments demonstrate state-of-the-art performance for both source separation and source localization, particularly in high levels of background noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- MESH2IR: Neural Acoustic Impulse Response Generator for Complex 3D ScenesAnton Ratnarajah, Zhenyu Tang, Rohith Aralikatti, Dinesh ManochaACM MM 2022 · 35 citations
- Look Once to Hear: Target Speech Hearing with Noisy ExamplesBandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka et al.CHI 2024 · 25 citations
- UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training MixturesZhong-Qiu Wang, Shinji WatanabeNeurIPS 2023 · 24 citations
- GWA: A Large High-Quality Acoustic Dataset for Audio ProcessingZhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, Dinesh ManochaSIGGRAPH 2022 · 23 citations
- Hybrid Neural Networks for On-Device Directional HearingAnran Wang, Maruchi Kim, Hao Zhang, Shyamnath GollakotaAAAI 2022 · 18 citations
Builds on2
Related papers
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 39 citations
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao et al.INFOCOM 2026 · 1 citation
- Filter-Recovery Network for Multi-Speaker Audio-Visual Speech SeparationHaoyue Cheng, Zhaoyang Liu, Wayne Wu, Limin WangICLR 2023
- Learning to Separate Voices by Spatial RegionsAlan Xu, Romit Roy ChoudhuryICML 2022 · 18 citations
- DeepRange: Acoustic Ranging via Deep LearningWenguang Mao, Wei Sun, Mei Wang, Lili QiuUbiComp 2021 · 49 citations
